STEP analysis evaluation

draftwright.evaluation.step_analysis measures STEP-analysis capability against a versioned, independently authored corpus. It is an engineering evaluation surface, not a drawing-lint score and not a replacement for per-drawing diagnostics.

Evidence boundary

The corpus JSON owns the denominator, physical identity fields, expected parameter values, tolerances, required downstream stages, fixture hashes, licence, and provenance. None is generated from RecognitionResult, feature_census(), a capability declaration, or the drawing under test. An observer may only normalize actual evidence; ObservedFact deliberately has no oracle/case fact identifier. Expected and observed facts are matched by family plus authored physical identity fields.

Format 1 currently proves independent holes, countersinks, hole-patterns, double-d-bores, flats, pockets, pocket-patterns, grooves, rectangular-pads, polygonal-bosses, plates, chamfers and fillets vertical slices. Each observer reads released quiddity geometry records, builds one drawing, and reads all four downstream outcomes from that build through the public IR, Sheet, generated-code and ADR 5 (was 0010) provenance seams.

Since #1217 that outcome comes from the engine's own requirement ledger (linting.hole_coverage.hole_requirement_outcomes) rather than from a second correspondence implementation here, and the ledger is treated as a pointer, not proof: supported requires both that the engine recorded an annotation as carrying the hole's size and that the annotation renders the value the compiler approved, checked through linting.evidence. unsupported means the requirement was not placed, or the annotation contradicts it. unknown means the ledger declined to join the hole to a feature without guessing — it scores as a miss, so it is an honest label rather than an exemption (#1202, #1206). Pattern arrangement dimensions follow the same ledger-pointer rule: placement provenance is insufficient unless the rendered pitch or BCD agrees with the compiler-approved value. Adding another family requires its own independently authored fixtures, facts, identity fields, observer, downstream evidence, and corpus version change; copying a manifest name alone cannot enlarge the denominator.

Scores and units

The evaluator reports three layers and never manufactures a composite:

  1. Detection uses a deterministic maximum bipartite match. Recall is matched expected physical facts divided by all expected physical facts. The reported false-positive rate is unmatched observations divided by all observations; when nothing was observed it is 0.0. Recall is unavailable (None) for a genuinely empty negative-case denominator.
  2. Parameter fidelity checks each authored parameter on a matched fact. A value receives one unit only when it is present and within the authored absolute tolerance. There is no tolerance interpolation: partial credit exists only across independently listed parameter units. The score is passed units divided by all checked units, or None when detection produced no matched parameter-bearing fact.
  3. Downstream usefulness checks each independently required IR-adapter, Sheet declaration, generated-code, and drawing-consumer boundary. supported receives one unit; missing, deferred, unsupported, and unknown states receive zero. The score is passed boundaries over required boundaries, or None when no matched fact has a downstream requirement.

Corpus aggregation is micro-averaged from raw units, not an average of case percentages, so many small cases cannot outweigh a missed compound case. Per-case diagnostics identify the layer, family, parameter/boundary, expected value, and observation. Inputs and diagnostics are sorted by canonical content, so neither expected-fact order, recogniser output order, nor STEP entity order changes a result.

unknown and unsupported are explicit analysis outcomes. They never count as complete. A corpus may independently expect one for an ambiguous or intentionally unsupported case, in which case the case is conformant—the system answered honestly—but still not complete. This distinction avoids rewarding fabricated certainty. A supported negative case with no expected and no observed facts can be complete within the corpus's stated family scope.

Corpus and determinism

The initial hole corpus lives in tests/fixtures/evaluation/corpus-v1.json; the independent hole-pattern arrangement corpus lives beside it as corpus-hole-patterns-v1.json. Each contains positive, negative, ambiguous, compound, and topology-order-variant cases. Every STEP file is hash-pinned and CC0-licensed, and its construction-derived oracle is documented beside it. Each topology pair has identical geometry and expected facts but bijectively renumbered, reverse-serialized Part 21 entities. CI runs the same evaluator on every supported Python version; repeated evaluation and each pair must produce identical layer results.

The anti-self-validation mutation replaces the real hole observer with an empty observer. Expected facts remain five, matches fall to zero, and recall falls from 1.0 to 0.0. Weakening or deleting a recogniser therefore cannot shrink the benchmark denominator and preserve a perfect result.

The pattern corpus owns one arrangement fact for each accepted group. Its identity contains the canonical member sites, while its scored parameters contain only the group grammar: count and BCD, linear pitch/direction, or grid rows/columns/pitches/angle/centre. It deliberately does not repeat member diameter, depth, bottom or individual-location requirements from the hole corpus. Provider patterns must reference the exact accepted aggregate HoleRecord members, and those member sets must be disjoint, so N:1 grouping cannot become a second physical-hole denominator.

The countersink corpus owns one conical seat per signed mouth axis and physical opening centre. It scores opening diameter, drill diameter, included angle, and geometric depth independently, while the existing hole ledger remains the sole owner of bore and finished-drawing requirements. A seat earns downstream credit only by following the exact provider-owned HoleRecord.csink object into both countersink.diameter and countersink.angle compiler identities and role-specific confirmed placed ink; it never rematches geometry or recounts the parent bore. Object identity names the provider-selected hole, while the public semantic predicate validates that ownership exactly once. Canonical-site collisions across disconnected coaxial bodies fail closed. Equal seats may share one grouped callout without collapsing their two physical observations. A second seat on the opposite face of one bore remains explicit but unverifiable at the downstream boundaries because the current HoleRecord and HoleFeature waist is singular; the completeness ledger counts both unverifiable dimensional outcomes rather than hiding the occurrence. The seven-case corpus covers a plain negative, the external-cone false-positive regression, a deburr ambiguity, a positive seat, a mixed-size ownership pair, and a reverse-serialized equal-seat pair.

The Double-D corpus owns one profiled through-bore occurrence per complete physical frame. Full axis location keeps disconnected coaxial bodies distinct; major diameter, A/F, through depth, through state and canonical flat direction are parameters rather than identity. Exact inventory multiplicity must survive automatic IR, Sheet.double_d_bore, generated code and one placed ⌀major THRU DOUBLE-D across A/F statement carrying both compiler identities. The provider aggregate's exclusive ownership keeps the parent circular void out of the ordinary-hole denominator.

The flat corpus owns one physical across-flats fact per stock axis line and connected axial span. Opposed provider faces on one Double-D body therefore group into one fact, while parallel lobes and disjoint coaxial bodies remain independent. Axis, canonical direction, axis line and stock span form identity; across-flats size, contributing face count and face anchors are scored parameters. The observer reads the build-owned recognition aggregate once and follows each group through the automatic IR, public Sheet.flat declaration, executed generated Sheet code and placed semantic measurement provenance.

The groove corpus owns one annular recess per shaft axis line and station, scoring axial width and floor diameter through one exact WIDE × ø statement. The chamfer corpus owns one planar or conical bevel per axis, physical anchor and surface form, scoring both legs and angle through exact C or leg × angle ink. The fillet corpus owns one cylindrical or toroidal round per axis, physical surface anchor and form, scoring radius through exact R or grouped n× R ink. All three verify a live physical leader target and include compound and topology-order controls; the chamfer corpus additionally pins AngledStep ownership, while the fillet corpus pins CircularBlindStep ownership.

The polygonal-boss corpus owns one attached regular hexagonal prism per principal axis and physical centre. It scores the provider's six-side schema invariant, across-flats, height and canonical physical flat-support pairs while keeping whole polygonal stock, recesses, detached prisms, circular bosses and rectangular pads outside the family denominator. Both A/F and height must survive through exact compiler identities. The finished A/F arrow is checked against a retained flat centre only after source-to-IR semantic correspondence is established, so rendered page geometry validates usefulness but never selects the feature owner.

The Plate corpus owns one thin material slab per body-local multi-plate prismatic occurrence. A single flat slab is envelope-owned and contributes no duplicate Plate fact; detached single slabs, thick block-scale spans and rotational bodies are negative controls. Axis, axial midpoint and both independently authored transverse witness coordinates form every occurrence identity; thickness is the scored parameter. Exact provider-to-IR correspondence retains the full axis, interval and both witness coordinates. The drawing boundary accepts either a compiler-confirmed, solver-placed thickness.length Dimension or the explicit derived opposite wall of a complete U-channel chain. Exact envelope, step-level/shoulder, slot-pattern and attached polygonal-boss ownership prevents derived material spans from inflating the Plate denominator and fails closed when full-witness body-local support cannot be established. Raw boss ownership additionally requires a valid support ring, a single-solid part, and a complete boss-plus-slab envelope span. Plural-solid inventories remain unverifiable because Plate carries no body provenance. Drawing credit for a derived span requires verified finished claims for every dependency. It never infers coverage from annotation names, labels, views or page coordinates. The 11 cases own 20 physical facts, 20 parameter units and 80 downstream units; deleting every provider Plate therefore leaves 20 misses rather than shrinking the denominator.

from pathlib import Path

from draftwright.evaluation.step_analysis import evaluate_step_corpus, load_corpus

corpus = load_corpus(Path("tests/fixtures/evaluation/corpus-v1.json"))
result = evaluate_step_corpus(corpus)
print(result.detection.recall)
print(result.parameter_fidelity.score)
print(result.downstream_usefulness.score)

Versioning and compatibility

Every result is meaningful only with both versions:

  • metric_version is an integer protocol version. Change it for matching, units, denominators, aggregation, partial-credit, or outcome semantics. Results across metric versions are not comparable.
  • corpus_version follows SemVer. Patch releases may correct prose/provenance without changing fixtures, facts, tolerances, scope, or scores. Minor releases may add cases or independently evaluated families and establish a new additive baseline. Major releases remove cases or change existing geometry, expected facts, identity fields, tolerances, or required downstream states.

A fixture hash change is never silent: it requires a corpus version decision and review of the construction oracle. CI and reports must retain both versions and per-layer raw counts. Thresholds must name an exact metric version and corpus compatibility range rather than comparing anonymous percentages.

Relationship to drawing quality

This metric asks whether STEP geometry was detected faithfully and can reach Draftwright's owned consumer boundaries. Drawing.lint() asks whether one concrete authored/generated drawing is semantically and visually acceptable. Its quality.completeness.audited_score is conditional on requirements the current recognition run already found, so it can diagnose downstream omission but cannot measure physical recall. The independent corpus can measure recall but cannot certify an arbitrary production part or a rendered sheet. Keep the feature census descriptive, use this evaluation for regression/coverage claims, and use lint codes plus separate drawing completeness, restraint, legibility, and fidelity components for a drawing decision (#1176 added fidelity: whether what the drawing says is true, which the other three do not ask).

API

step_analysis

Evidence-based STEP analysis evaluation (#1169).

This module scores recogniser observations against an independently authored oracle. Its expectations do not inspect RecognitionResult or the feature census: adapters supply observations, while the benchmark case supplies the denominator and tolerances.

Every downstream boundary is OBSERVED through its real seam (#1369), never copied from the capability declaration: the built PartModel for ir_adapter; an explicit public Sheet.hole declaration for dsl_declaration; an executed emit_sheet_script result for generated_code; and the placed drawing's ADR 5 (was 0010) measurement provenance for drawing_consumer. The existing hole-requirement ledger supplies one conservative recognition-to-IR correspondence implementation for all four observations. It is a join, not the benchmark denominator: the independently authored corpus remains the only source of expected facts.

The hole-pattern slice (#1370) uses the same four boundaries through Sheet.pattern and the existing hole-requirement correspondence. Its separate corpus scores one arrangement fact per aggregate pattern. Member diameter/depth/bottom/location requirements stay solely in the hole corpus, so the derived N:1 group never becomes a second physical-hole denominator.

The Double-D slice (#1370) counts one profiled through-bore occurrence per full physical frame. Major diameter, A/F, through-depth and the unoriented flat line are scored independently, while automatic IR, public Sheet.double_d_bore, executed generated code and exact role-specific ⌀… DOUBLE-D … A/F ink must all retain the same occurrence. This is a separate profile fact, not a second ordinary-hole fact: provider aggregate ownership excludes its circular parent from RecognitionResult.holes.

The flat slice (#1371) scores one physical A/F requirement per stock line and axial span. Two opposed faces on one Double-D are member evidence for one requirement; equal parallel stock and disjoint coaxial stock remain separate facts. Across-flats and the face anchors are parameters, not benchmark identity, so weakening either lowers fidelity instead of hiding as a detection mismatch.

The lone-pocket slice (#1372) excludes members owned by pocket patterns and counts width, length, depth, plus two independently observed datum-location axes for an interior recess. Edge-anchored corner interruptions retain their three explicit sizes while their position is intentionally implicit. Opening side remains identity so opposed-face pockets cannot collapse at the IR waist.

The pocket-pattern slice (#1372) counts one grouped physical arrangement, never its member pockets again. Its width, length, depth, count, lattice and centre are observed through the same four boundaries, including exact count/pitch/location ink backed by compiler provenance.

The groove slice (#1372) counts one annular recess at each turning-axis station. Axis and physical anchor identify the occurrence; axial width and floor diameter are scored parameters and required drawing measurements. The drawing observation requires both compiler-approved identities on one exact semantic WIDE × ø callout, while the corpus remains independent of recognition output.

The rectangular-pad slice (#1372) counts one bounded protrusion at each signed attachment-plane centre. Principal axis, material-outward direction, and attachment point identify the occurrence; footprint width/length and local height are scored parameters. The drawing observation follows all five physical requirements through compiler measurement identities and structured directional location facts without treating the older geometric coverage fallback as semantic evidence.

The Plate slice (#1373) counts one body-local thin slab whose thickness is not already owned by the whole-part envelope. Thin axis, physical slab station and both independently authored transverse witness coordinates identify the occurrence; thickness is the scored parameter. Automatic IR, public declaration and executed generated code must preserve the full witness, and the drawing must carry the exact compiler-owned thickness identity and verified ink.

The polygonal-boss slice (#1372) counts one attached regular prism per principal axis and physical centre. Side count, A/F, height, ordered flat directions and physical flat centres are scored parameters. Automatic IR, public declaration and executed generated code must retain the complete prism record; the drawing must carry the exact A/F and height compiler identities, valid statement ink and a live A/F leader on one retained physical support face.

The polygonal-stock slice (#1371) counts one complete regular-hexagonal-prism body. Principal axis and physical centre identify the occurrence; side count, A/F, axial length and the coupled ring of flat directions/physical centres are scored parameters. Exact cap span and support geometry remain load-bearing correspondence evidence through automatic IR, public declaration and generated code. The drawing must carry both compiler identities, exact statement ink, and a live A/F leader on one retained physical support face. Attached bosses, machined/irregular prisms and compounds remain outside this whole-stock denominator.

The chamfer slice (#1374) counts one planar or conical bevel per physical anchor. Axis, anchor and surface form identify the occurrence; both legs and angle are scored parameters. One compiler identity reaches exact C or leg × angle ink and a live leader on the bevel/profile station; equal specifications may share ink only while retaining every member identity.

The fillet slice (#1374) counts one cylindrical or toroidal round per physical surface anchor. Axis, anchor and planar/turned form identify the occurrence; radius is the scored parameter. One compiler identity per round reaches exact R or grouped n× R ink and a live leader on the round/profile station. Aggregate ownership excludes curved walls assigned to CircularBlindStep.

The turned-step slice (#1374) counts one outside-diameter band per body-local axis line and axial station. Length and diameter are independently scored parameters and drawing requirements. A band uniquely owned by a correlated groove remains solely in the groove denominator; ambiguous nested/coaxial groove ownership is refused under the provider contract rather than guessed.

Known limit of the drawing observation: it reads the ADR 5 (was 0010) provenance seam, which registry.measurement_of carries and which is populated one render pass at a time (the set of tagged renderers is enumerated by tests/test_audit_differential.py, not by prose here — that docstring warns the prose version was wrong when first written). An un-tagged render pass therefore reads as a genuine omission, and this is a CLASS of limitation rather than a single case. Two instances are known:

  • the hole-table escalation, which withdraws the individual callouts and records the substitution on the table — admitted here via that ledger;
  • a turned part, where the bore's diameter reaches the sheet as a Leader but the hole requirement ledger still reports the bore size as missing. The benchmark therefore reports a loss for a hole whose size is visibly printed. No corpus fixture is turned today; adding one without closing that correspondence gap would make the number wrong.

A new representation route must be admitted here or it registers as a false loss.

Every draftwright import in this module is deliberately inside a function body — there are no module-level ones at all — which is the #313 lazy-load pattern rather than an accident. (quiddity counts: importing it puts build123d in sys.modules, so it carries the same cost.) It is load-bearing: importing this module costs ~0.01 s, and hoisting ANY engine import makes it one to two seconds, because every one pulls build123d transitively. Measured in a single process, the cost is essentially all build123d and is paid once — the draftwright modules themselves are free once it is loaded::

build123d                          (the whole cost)
draftwright.linting.hole_coverage  ~0.02-0.05 s   (after build123d)
draftwright.model.compiled         ~0.01-0.03 s
draftwright.linting.evidence        0.000 s
draftwright.builder                ~0.01-0.02 s

Absolute seconds are deliberately not quoted for build123d: measurements on two machines gave 1.35 s and 2.26 s. The SHAPE is the point and it reproduces. (An earlier version listed four figures of 1.4-2.0 s, one per module, from four separate cold processes — the same one-time cost measured four times and presented as if the modules differed. They do not.)

1229 filed three of these imports as "unexplained, hoist or justify"; measuring is what showed

the filing was wrong, and this note is the justification it asked for. Keep new engine imports inside the bodies too.

CorpusError

Bases: ValueError

The independent benchmark corpus is malformed or its evidence changed.

ObservationError

Bases: RuntimeError

A family observer could not distinguish an honest empty inventory from failure.

ParameterExpectation dataclass

An independently authored value and its absolute acceptance tolerance.

ExpectedFact dataclass

One physical fact in the benchmark denominator.

ObservedFact dataclass

One normalized fact emitted by the system under evaluation.

No benchmark identifier is accepted. Matching is derived from family and independently specified physical identity fields, preventing an adapter from copying the oracle's answer.

BenchmarkCase dataclass

One independently sourced STEP fixture and its expected facts.

BenchmarkCorpus dataclass

A validated, versioned collection of independently authored cases.

CaseEvaluation dataclass

conformant property

Whether the observation matches the oracle, including an honest non-answer.

complete property

Whether every independently expected layer is satisfied without false claims.

CorpusEvaluation dataclass

Micro-averaged evidence layers; deliberately no composite scalar.

evaluate_case(case, *, observations, outcome='supported')

Score one case without deriving either expectations or tolerances from observations.

evaluate_corpus(cases, *, corpus_version=None)

Aggregate raw units across cases without averaging away small-case failures.

load_corpus(path)

Load and fail-closed validate a corpus, including every fixture hash.

evaluate_step_corpus(corpus, *, observers=None, outcomes=None)

Import every pinned STEP fixture and evaluate normalized family observations.