Skills Assessment¶
Formal competence assessment (diagnostic) — instantiates Competence Calibration Feedback
A formal, domain-scoped evaluation that scores actual performance against an explicit standard, producing the evidence a calibration loop compares self-assessment to.
Skills Assessment is the formal, graded measure that produces hard evidence of what a person can actually do. Its defining move is standardization against an explicit standard: a structured evaluation, bounded to a named competence, scored against a defined criterion, yielding a competence reading that is independent of how the person feels about their ability. It is the objective side of the loop — and only that side. On its own it is evidence, not calibration; it becomes Competence Calibration Feedback only when its result is set against a self-assessment and used to update confidence, learning, or authority. Where a Simulation or Case Test manufactures a scenario to probe transfer, a skills assessment applies a standardized, scored measure to certify a level.
Example¶
A structural welder can't self-certify their way onto a bridge project; they sit a qualification test. The assessment is tightly scoped — a specific joint configuration, position, and material — performed under controlled conditions, and the resulting weld is scored against an explicit code standard, including destructive and non-destructive examination of the sample.[1] The output is not the welder's opinion of their skill and not a vibe from a supervisor; it is a pass-against-standard on a named, bounded competence. That reading is exactly the evidence the calibration loop needs: laid against the welder's own confidence, it either confirms it or corrects it, and it feeds the decision about what work they may be assigned. The assessment itself just supplies the hard number; the calibration is what someone does with it next.
How it works¶
Its distinguishing features are standardized tasks and explicit, defined scoring, bounded to a specific domain. Standardization makes results comparable across people and over time; scoring against a named standard makes the reading anchored rather than relative to whoever happened to judge. It deliberately measures the evidence side and stops — it captures no self-assessment and delivers no feedback. Its whole job is to convert "can they actually do it?" into a defensible, comparable measurement.
Tuning parameters¶
- Standardization — fixed tasks versus adaptive ones. Fixed enables clean comparison; adaptive targets the individual's edge but sacrifices comparability.
- Validity focus — whether the assessment measures the real competence or a convenient proxy (a knowledge test standing in for a performance test). Performance measures are more valid but costlier.
- Scoring objectivity — rubric-scored versus expert judgment. Objective scoring is comparable and defensible; expert judgment catches nuance a rubric misses.
- Stakes — low-stakes diagnostic versus high-stakes certification. High stakes motivate effort but invite gaming, coaching-to-the-test, and anxiety distortion.
- Domain fit — how tightly the assessment's tasks sample the actual work the competence is claimed for.
When it helps, and when it misleads¶
Its strength is supplying the hard, comparable evidence that makes self-assessment corrigible: without a real measure of what someone can do, calibration is just competing opinions.
It misleads because an assessment only measures what it samples. A valid-looking test of a proxy skill will confidently mis-certify — the score is precise but about the wrong thing, a failure of construct validity.[2] High-stakes tests get gamed and coached-to, and a score can harden into an identity label ("I'm a 3") rather than staying a data point. And a score sitting unused beside no self-assessment calibrates nothing at all. The discipline is to validate that the test actually samples the real competence, keep it fair, and feed the result back into the comparison with self-assessment rather than filing it as a grade.
How it implements the components¶
Skills Assessment fills the domain-scoped, standardized-evidence slice:
competence_domain— scopes the assessment to a bounded, named capability, so the result is about a specific competence rather than general ability.performance_evidence_set— its scored result is objective, domain-relevant evidence of actual performance.performance_benchmark— scores against an explicit, defined standard, so the reading is anchored rather than relative to the grader.
It does not capture the self-assessment its result will be compared with (self_assessment, Confidence Rating Scale) or explain the resulting gap (calibration_feedback, Benchmarked Feedback).
Related¶
- Instantiates: Competence Calibration Feedback — supplies the standardized, scored evidence the loop compares self-assessment against.
- Sibling mechanisms: Simulation or Case Test · Competency Framework · Benchmarked Feedback · Calibration Conversation · Confidence Rating Scale · Calibration Exercise · Exemplar Comparison · Peer Review · Reflective Error Log · Decision Rights by Competence · Supervised Practice
Notes¶
A skills assessment is evidence, not calibration. A drawer full of scores that are never set against people's self-assessments or used to adjust learning and authority changes nothing — the assessment implements the archetype only as the input to a comparison, not as an end in itself.
References¶
[1] Welder qualification testing against a structural welding code (such as AWS D1.1) is a real, standardized skills assessment: a welder produces a specified test weld under defined conditions, and the sample is examined and scored against explicit acceptance criteria. It certifies a bounded competence, not general ability. ↩
[2] Construct validity is the degree to which a test actually measures the underlying construct it claims to — here, real competence rather than a proxy. An assessment can be perfectly reliable and still invalid if it scores the wrong thing, which is why sampling the real domain of work matters more than the tidiness of the score. ↩