STEP analysis evaluation¶
draftwright.evaluation.step_analysis measures STEP-analysis capability against a versioned,
independently authored corpus. It is an engineering evaluation surface, not a drawing-lint score
and not a replacement for per-drawing diagnostics.
Evidence boundary¶
The corpus JSON owns the denominator, physical identity fields, expected parameter values,
tolerances, required downstream stages, fixture hashes, licence, and provenance. None is generated
from RecognitionResult, feature_census(), a capability declaration, or the drawing under test.
An observer may only normalize actual evidence; ObservedFact deliberately has no oracle/case fact
identifier. Expected and observed facts are matched by family plus authored physical identity fields.
Format 1 currently proves independent holes, countersinks, hole-patterns,
double-d-bores, flats,
pockets, pocket-patterns, grooves, rectangular-pads, polygonal-bosses, plates,
chamfers and fillets vertical slices. Each observer reads released quiddity
geometry records, builds one drawing, and reads all four downstream outcomes from that build
through the public IR, Sheet, generated-code and ADR 5 (was 0010) provenance seams.
Since #1217 that outcome comes from the engine's own requirement ledger
(linting.hole_coverage.hole_requirement_outcomes) rather than from a second correspondence
implementation here, and the ledger is treated as a pointer, not proof: supported requires
both that the engine recorded an annotation as carrying the hole's size and that the annotation
renders the value the compiler approved, checked through linting.evidence. unsupported means
the requirement was not placed, or the annotation contradicts it. unknown means the ledger
declined to join the hole to a feature without guessing — it scores as a miss, so it is an honest
label rather than an exemption (#1202, #1206). Pattern arrangement dimensions follow the same
ledger-pointer rule: placement provenance is insufficient unless the rendered pitch or BCD agrees
with the compiler-approved value. Adding another family requires its own independently authored
fixtures, facts, identity fields, observer, downstream evidence, and corpus version change; copying
a manifest name alone cannot enlarge the denominator.
Scores and units¶
The evaluator reports three layers and never manufactures a composite:
- Detection uses a deterministic maximum bipartite match. Recall is matched expected physical
facts divided by all expected physical facts. The reported false-positive rate is unmatched
observations divided by all observations; when nothing was observed it is
0.0. Recall is unavailable (None) for a genuinely empty negative-case denominator. - Parameter fidelity checks each authored parameter on a matched fact. A value receives one
unit only when it is present and within the authored absolute tolerance. There is no tolerance
interpolation: partial credit exists only across independently listed parameter units. The
score is passed units divided by all checked units, or
Nonewhen detection produced no matched parameter-bearing fact. - Downstream usefulness checks each independently required IR-adapter,
Sheetdeclaration, generated-code, and drawing-consumer boundary.supportedreceives one unit; missing,deferred,unsupported, and unknown states receive zero. The score is passed boundaries over required boundaries, orNonewhen no matched fact has a downstream requirement.
Corpus aggregation is micro-averaged from raw units, not an average of case percentages, so many small cases cannot outweigh a missed compound case. Per-case diagnostics identify the layer, family, parameter/boundary, expected value, and observation. Inputs and diagnostics are sorted by canonical content, so neither expected-fact order, recogniser output order, nor STEP entity order changes a result.
unknown and unsupported are explicit analysis outcomes. They never count as complete. A corpus
may independently expect one for an ambiguous or intentionally unsupported case, in which case the
case is conformant—the system answered honestly—but still not complete. This distinction avoids
rewarding fabricated certainty. A supported negative case with no expected and no observed facts
can be complete within the corpus's stated family scope.
Corpus and determinism¶
The initial hole corpus lives in tests/fixtures/evaluation/corpus-v1.json; the independent
hole-pattern arrangement corpus lives beside it as corpus-hole-patterns-v1.json. Each contains
positive, negative, ambiguous, compound, and topology-order-variant cases. Every STEP file is
hash-pinned and CC0-licensed, and its construction-derived oracle is documented beside it. Each
topology pair has identical geometry and expected facts but bijectively renumbered,
reverse-serialized Part 21 entities. CI runs the same evaluator on every supported Python version;
repeated evaluation and each pair must produce identical layer results.
The anti-self-validation mutation replaces the real hole observer with an empty observer. Expected
facts remain five, matches fall to zero, and recall falls from 1.0 to 0.0. Weakening or deleting
a recogniser therefore cannot shrink the benchmark denominator and preserve a perfect result.
The pattern corpus owns one arrangement fact for each accepted group. Its identity contains the
canonical member sites, while its scored parameters contain only the group grammar: count and BCD,
linear pitch/direction, or grid rows/columns/pitches/angle/centre. It deliberately does not repeat
member diameter, depth, bottom or individual-location requirements from the hole corpus. Provider
patterns must reference the exact accepted aggregate HoleRecord members, and those member sets
must be disjoint, so N:1 grouping cannot become a second physical-hole denominator.
The countersink corpus owns one conical seat per signed mouth axis and physical opening centre. It
scores opening diameter, drill diameter, included angle, and geometric depth independently, while
the existing hole ledger remains the sole owner of bore and finished-drawing requirements. A seat
earns downstream credit only by following the exact provider-owned HoleRecord.csink object into
both countersink.diameter and countersink.angle compiler identities and role-specific confirmed
placed ink; it never rematches geometry or recounts the parent bore. Object identity names the
provider-selected hole, while the public semantic predicate validates that ownership exactly once.
Canonical-site collisions across disconnected coaxial bodies fail closed. Equal seats may share one
grouped callout without collapsing their two physical observations. A second seat on the opposite face of one bore
remains explicit but unverifiable at the downstream boundaries because the current HoleRecord and
HoleFeature waist is singular; the completeness ledger counts both unverifiable dimensional
outcomes rather than hiding the occurrence. The seven-case corpus covers a plain negative, the
external-cone false-positive regression, a deburr ambiguity, a positive seat, a mixed-size ownership
pair, and a reverse-serialized equal-seat pair.
The Double-D corpus owns one profiled through-bore occurrence per complete physical frame. Full
axis location keeps disconnected coaxial bodies distinct; major diameter, A/F, through depth,
through state and canonical flat direction are parameters rather than identity. Exact inventory
multiplicity must survive automatic IR, Sheet.double_d_bore, generated code and one placed
⌀major THRU DOUBLE-D across A/F statement carrying both compiler identities. The provider
aggregate's exclusive ownership keeps the parent circular void out of the ordinary-hole
denominator.
The flat corpus owns one physical across-flats fact per stock axis line and connected axial span.
Opposed provider faces on one Double-D body therefore group into one fact, while parallel lobes and
disjoint coaxial bodies remain independent. Axis, canonical direction, axis line and stock span
form identity; across-flats size, contributing face count and face anchors are scored parameters.
The observer reads the build-owned recognition aggregate once and follows each group through the
automatic IR, public Sheet.flat declaration, executed generated Sheet code and placed semantic
measurement provenance.
The groove corpus owns one annular recess per shaft axis line and station, scoring axial width and
floor diameter through one exact WIDE × ø statement. The chamfer corpus owns one planar or
conical bevel per axis, physical anchor and surface form, scoring both legs and angle through exact
C or leg × angle ink. The fillet corpus owns one cylindrical or toroidal round per axis, physical
surface anchor and form, scoring radius through exact R or grouped n× R ink. All three verify a
live physical leader target and include compound and topology-order controls; the chamfer corpus
additionally pins AngledStep ownership, while the fillet corpus pins CircularBlindStep ownership.
The polygonal-boss corpus owns one attached regular hexagonal prism per principal axis and physical centre. It scores the provider's six-side schema invariant, across-flats, height and canonical physical flat-support pairs while keeping whole polygonal stock, recesses, detached prisms, circular bosses and rectangular pads outside the family denominator. Both A/F and height must survive through exact compiler identities. The finished A/F arrow is checked against a retained flat centre only after source-to-IR semantic correspondence is established, so rendered page geometry validates usefulness but never selects the feature owner.
The Plate corpus owns one thin material slab per body-local multi-plate prismatic occurrence. A
single flat slab is envelope-owned and contributes no duplicate Plate fact; detached single slabs,
thick block-scale spans and rotational bodies are negative controls. Axis, axial midpoint and both
independently authored transverse witness coordinates form every occurrence identity; thickness is
the scored parameter. Exact provider-to-IR correspondence retains
the full axis, interval and both witness coordinates. The drawing boundary accepts either a
compiler-confirmed, solver-placed thickness.length Dimension or the explicit derived opposite
wall of a complete U-channel chain. Exact envelope, step-level/shoulder, slot-pattern and attached
polygonal-boss ownership prevents derived material spans from inflating the Plate denominator and
fails closed when full-witness body-local support cannot be established. Raw boss ownership
additionally requires a valid support ring, a single-solid part, and a complete boss-plus-slab
envelope span. Plural-solid inventories remain unverifiable because Plate carries no body
provenance. Drawing credit for a
derived span requires verified finished claims for every dependency. It never infers coverage from annotation
names, labels, views or page coordinates. The 11 cases own 20 physical facts, 20 parameter units and 80 downstream
units; deleting every provider Plate therefore leaves 20 misses rather than shrinking the
denominator.
from pathlib import Path
from draftwright.evaluation.step_analysis import evaluate_step_corpus, load_corpus
corpus = load_corpus(Path("tests/fixtures/evaluation/corpus-v1.json"))
result = evaluate_step_corpus(corpus)
print(result.detection.recall)
print(result.parameter_fidelity.score)
print(result.downstream_usefulness.score)
Versioning and compatibility¶
Every result is meaningful only with both versions:
metric_versionis an integer protocol version. Change it for matching, units, denominators, aggregation, partial-credit, or outcome semantics. Results across metric versions are not comparable.corpus_versionfollows SemVer. Patch releases may correct prose/provenance without changing fixtures, facts, tolerances, scope, or scores. Minor releases may add cases or independently evaluated families and establish a new additive baseline. Major releases remove cases or change existing geometry, expected facts, identity fields, tolerances, or required downstream states.
A fixture hash change is never silent: it requires a corpus version decision and review of the construction oracle. CI and reports must retain both versions and per-layer raw counts. Thresholds must name an exact metric version and corpus compatibility range rather than comparing anonymous percentages.
Relationship to drawing quality¶
This metric asks whether STEP geometry was detected faithfully and can reach Draftwright's owned
consumer boundaries. Drawing.lint() asks whether one concrete authored/generated drawing is
semantically and visually acceptable. Its quality.completeness.audited_score is conditional on
requirements the current recognition run already found, so it can diagnose downstream omission but
cannot measure physical recall. The independent corpus can measure recall but cannot certify an
arbitrary production part or a rendered sheet. Keep the feature census descriptive, use this
evaluation for regression/coverage claims, and use lint codes plus separate drawing completeness,
restraint, legibility, and fidelity components for a drawing decision (#1176 added
fidelity: whether what the drawing says is true, which the other three do not ask).
API¶
step_analysis
¶
Evidence-based STEP analysis evaluation (#1169).
This module scores recogniser observations against an independently authored oracle. Its
expectations do not inspect RecognitionResult or the feature census: adapters supply
observations, while the benchmark case supplies the denominator and tolerances.
Every downstream boundary is OBSERVED through its real seam (#1369), never copied from the
capability declaration: the built PartModel for ir_adapter; an explicit public
Sheet.hole declaration for dsl_declaration; an executed emit_sheet_script result
for generated_code; and the placed drawing's ADR 5 (was 0010) measurement provenance for
drawing_consumer. The existing hole-requirement ledger supplies one conservative
recognition-to-IR correspondence implementation for all four observations. It is a join, not
the benchmark denominator: the independently authored corpus remains the only source of
expected facts.
The hole-pattern slice (#1370) uses the same four boundaries through Sheet.pattern and the
existing hole-requirement correspondence. Its separate corpus scores one arrangement fact per
aggregate pattern. Member diameter/depth/bottom/location requirements stay solely in the hole
corpus, so the derived N:1 group never becomes a second physical-hole denominator.
The Double-D slice (#1370) counts one profiled through-bore occurrence per full physical frame.
Major diameter, A/F, through-depth and the unoriented flat line are scored independently, while
automatic IR, public Sheet.double_d_bore, executed generated code and exact role-specific
⌀… DOUBLE-D … A/F ink must all retain the same occurrence. This is a separate profile
fact, not a second ordinary-hole fact: provider aggregate ownership excludes its circular parent
from RecognitionResult.holes.
The flat slice (#1371) scores one physical A/F requirement per stock line and axial span. Two opposed faces on one Double-D are member evidence for one requirement; equal parallel stock and disjoint coaxial stock remain separate facts. Across-flats and the face anchors are parameters, not benchmark identity, so weakening either lowers fidelity instead of hiding as a detection mismatch.
The lone-pocket slice (#1372) excludes members owned by pocket patterns and counts width, length, depth, plus two independently observed datum-location axes for an interior recess. Edge-anchored corner interruptions retain their three explicit sizes while their position is intentionally implicit. Opening side remains identity so opposed-face pockets cannot collapse at the IR waist.
The pocket-pattern slice (#1372) counts one grouped physical arrangement, never its member pockets again. Its width, length, depth, count, lattice and centre are observed through the same four boundaries, including exact count/pitch/location ink backed by compiler provenance.
The groove slice (#1372) counts one annular recess at each turning-axis station. Axis and physical
anchor identify the occurrence; axial width and floor diameter are scored parameters and required
drawing measurements. The drawing observation requires both compiler-approved identities on one
exact semantic WIDE × ø callout, while the corpus remains independent of recognition output.
The rectangular-pad slice (#1372) counts one bounded protrusion at each signed attachment-plane centre. Principal axis, material-outward direction, and attachment point identify the occurrence; footprint width/length and local height are scored parameters. The drawing observation follows all five physical requirements through compiler measurement identities and structured directional location facts without treating the older geometric coverage fallback as semantic evidence.
The Plate slice (#1373) counts one body-local thin slab whose thickness is not already owned by the whole-part envelope. Thin axis, physical slab station and both independently authored transverse witness coordinates identify the occurrence; thickness is the scored parameter. Automatic IR, public declaration and executed generated code must preserve the full witness, and the drawing must carry the exact compiler-owned thickness identity and verified ink.
The polygonal-boss slice (#1372) counts one attached regular prism per principal axis and physical centre. Side count, A/F, height, ordered flat directions and physical flat centres are scored parameters. Automatic IR, public declaration and executed generated code must retain the complete prism record; the drawing must carry the exact A/F and height compiler identities, valid statement ink and a live A/F leader on one retained physical support face.
The polygonal-stock slice (#1371) counts one complete regular-hexagonal-prism body. Principal axis and physical centre identify the occurrence; side count, A/F, axial length and the coupled ring of flat directions/physical centres are scored parameters. Exact cap span and support geometry remain load-bearing correspondence evidence through automatic IR, public declaration and generated code. The drawing must carry both compiler identities, exact statement ink, and a live A/F leader on one retained physical support face. Attached bosses, machined/irregular prisms and compounds remain outside this whole-stock denominator.
The chamfer slice (#1374) counts one planar or conical bevel per physical anchor. Axis, anchor and
surface form identify the occurrence; both legs and angle are scored parameters. One compiler
identity reaches exact C or leg × angle ink and a live leader on the bevel/profile station;
equal specifications may share ink only while retaining every member identity.
The fillet slice (#1374) counts one cylindrical or toroidal round per physical surface anchor.
Axis, anchor and planar/turned form identify the occurrence; radius is the scored parameter. One
compiler identity per round reaches exact R or grouped n× R ink and a live leader on the
round/profile station. Aggregate ownership excludes curved walls assigned to CircularBlindStep.
The turned-step slice (#1374) counts one outside-diameter band per body-local axis line and axial station. Length and diameter are independently scored parameters and drawing requirements. A band uniquely owned by a correlated groove remains solely in the groove denominator; ambiguous nested/coaxial groove ownership is refused under the provider contract rather than guessed.
Known limit of the drawing observation: it reads the ADR 5 (was 0010) provenance seam, which
registry.measurement_of carries and which is populated one render pass at a time (the set
of tagged renderers is enumerated by tests/test_audit_differential.py, not by prose here —
that docstring warns the prose version was wrong when first written). An un-tagged render pass
therefore reads as a genuine omission, and this is a CLASS of limitation rather than a single
case. Two instances are known:
- the hole-table escalation, which withdraws the individual callouts and records the substitution on the table — admitted here via that ledger;
- a turned part, where the bore's diameter reaches the sheet as a
Leaderbut the hole requirement ledger still reports the bore size as missing. The benchmark therefore reports a loss for a hole whose size is visibly printed. No corpus fixture is turned today; adding one without closing that correspondence gap would make the number wrong.
A new representation route must be admitted here or it registers as a false loss.
Every draftwright import in this module is deliberately inside a function body — there are
no module-level ones at all — which is the #313 lazy-load pattern rather than an accident.
(quiddity counts: importing it puts build123d in sys.modules, so it carries the
same cost.) It is load-bearing: importing this module costs ~0.01 s, and hoisting ANY engine import makes it
one to two seconds, because every one pulls build123d transitively. Measured in a single process,
the cost is essentially all build123d and is paid once — the draftwright modules themselves are
free once it is loaded::
build123d (the whole cost)
draftwright.linting.hole_coverage ~0.02-0.05 s (after build123d)
draftwright.model.compiled ~0.01-0.03 s
draftwright.linting.evidence 0.000 s
draftwright.builder ~0.01-0.02 s
Absolute seconds are deliberately not quoted for build123d: measurements on two machines gave 1.35 s and 2.26 s. The SHAPE is the point and it reproduces. (An earlier version listed four figures of 1.4-2.0 s, one per module, from four separate cold processes — the same one-time cost measured four times and presented as if the modules differed. They do not.)
1229 filed three of these imports as "unexplained, hoist or justify"; measuring is what showed¶
the filing was wrong, and this note is the justification it asked for. Keep new engine imports inside the bodies too.
CorpusError
¶
Bases: ValueError
The independent benchmark corpus is malformed or its evidence changed.
ObservationError
¶
Bases: RuntimeError
A family observer could not distinguish an honest empty inventory from failure.
ParameterExpectation
dataclass
¶
An independently authored value and its absolute acceptance tolerance.
ExpectedFact
dataclass
¶
One physical fact in the benchmark denominator.
ObservedFact
dataclass
¶
One normalized fact emitted by the system under evaluation.
No benchmark identifier is accepted. Matching is derived from family and independently specified physical identity fields, preventing an adapter from copying the oracle's answer.
BenchmarkCase
dataclass
¶
One independently sourced STEP fixture and its expected facts.
BenchmarkCorpus
dataclass
¶
A validated, versioned collection of independently authored cases.
CaseEvaluation
dataclass
¶
CorpusEvaluation
dataclass
¶
Micro-averaged evidence layers; deliberately no composite scalar.
evaluate_case(case, *, observations, outcome='supported')
¶
Score one case without deriving either expectations or tolerances from observations.
evaluate_corpus(cases, *, corpus_version=None)
¶
Aggregate raw units across cases without averaging away small-case failures.
load_corpus(path)
¶
Load and fail-closed validate a corpus, including every fixture hash.
evaluate_step_corpus(corpus, *, observers=None, outcomes=None)
¶
Import every pinned STEP fixture and evaluate normalized family observations.