Principles · September 7, 2026 · 8 min
We found a fix that reached 76 more sections and did not merge it
A wider band window reached 162 sections, but the 76 new ones score median 1.106 ft against 0.323 ft on the same set. We did not merge it, and here is why.
We found a change to our own code that reached 76 more cross-sections off one plan set, and we did not merge it. Then we found a second change that worked — one section to forty out of fifty-six — and did not merge that one either. Both would have improved the counts a buyer is able to see. A measurement stopped the first one and a registered control stopped the second. Here are both, and why we think a shop that ships either one looks better than it is.
Mathyra is in private development. Nothing here was sold to anyone, and we have never shipped a surface.
The fix that reached 162 sections
Before our reader can digitize a cross-section it has to find the strip of the page the section is drawn in, and inside that strip, the horizontal scale — the thing that tells it where offset zero sits and how many feet an inch across the page is worth. On one 56-page cross-section volume, the reader lost that scale on 100 of 211 offset rulers. No scale, no section.
Widening the window the reader searches gets most of them back. With the wider window the fix reaches 162 sections. But the sections that come in new do not behave like the sections that were already there. The 76 new sections score a median of 1.106 ft. The 84 scored sections that came through both the old way and the new score a median of 0.323 ft. Both medians are scored on the same 56-page cross-section volume against the designer's grade line — the finished-grade surface in the LandXML that shipped in the same bid package as the sheets. That reference matters as much as the numbers. Cross-section sheets are plotted from the corridor design and the bid-package model is exported from it, so scoring one against the other measures how faithfully a drawing was plotted and read back, not how close either is to the ground. It is a fidelity comparison, and we use it as one. Those two groups account for 160 of the 162 sections the wider window reaches; we have not accounted for the other two here, and that gap is ours, not the sheet's.
A median near a foot is not a rounding difference. But it is a disagreement with another drawing-derived reference, not a measured distance from built ground — external accuracy is unmeasured, and this number does not change that.
We do not know why the new sections score worse — nothing in the repo measures the cause, and the band window is our code, not the drafter's. What we have is the split itself: the sections the old window missed score roughly three times worse than the ones it already found, and that split is what stopped the merge.
So the trade on offer was a much larger section count on that set, paid for with a much worse median on the sections that came in new. We did not take it.
The band window is not fixed. The code at HEAD still picks the scale the old way, and a test in the repo pins that deferral in place. The honest claim is not that we solved it. The honest claim is that the loss is now visible, counted, and attached to a number that says what turning it on would cost.
The fix that overturned a control
The second one is smaller and cleaner, and it is the better illustration.
We had a set where the reader was pulling exactly one stationed section out of fifty-six. A change to how sections get assigned took that from 1 to 40 of 56. It worked. It was not merged, because it overturned a registered control — a case frozen earlier with an answer already checked by hand. When a change makes a previously correct case wrong, you do not have a fix. You have a change that trades one error for another and happens to be pointed at the errors you were counting that week.
A later, different change reached 1 to 41 of 56 with every control still intact. The insight behind it was about the drawing, not about our reader: the drafter centres the station label rather than hanging it off one end. That one was merged.
Same set, same problem, better number, no control broken. The difference between the two is not the score. It is that the second one explains something true about how the sheet was drawn, and the first one only accommodated our reader.
| Change | What it buys | What it costs | Status |
|---|---|---|---|
| Wider band window | Reaches 162 sections on a 56-page cross-section volume that had lost the scale on 100 of 211 offset rulers | The 76 new sections score median 1.106 ft against 0.323 ft on the 84 carried-over sections, both against the designer's grade line | Deferred. A test at HEAD pins the deferral |
| First assignment change | 1 to 40 of 56 stationed sections | Overturns a registered control | Not merged |
| Label-centre change | 1 to 41 of 56 stationed sections | No control broken | Merged |
Why shipping either one makes a shop look better than it is
Take the numbers a buyer can actually inspect on a data-prep job: sections delivered and pages processed. Both of these changes move both the right way. We have not measured what either one does to runtime, so we are not going to claim it. Neither one moves anything a buyer can check in the other direction, because there is nothing in the package that measures error. There is no field shot in it. There is no independent check.
Our own output prints that limit rather than hiding it. The disclaimer, verbatim: "Not one coordinate in this deliverable has been compared against a surveyed value, a designer's model, or a hand-read elevation off the sheet."
And the obvious repair — score the sheets against the designer's model that came with the bid package — retires exactly one of those three legs, and not the one that matters most. Agencies document a workflow in which the cross-section sheets and the bid-package model come out of the same corridor design. We have not run a measurement behind that; we took it from how the agencies describe their own production, and we treat it as a limit on our numbers rather than as a finding of ours. We have adopted that as a binding limit on what our own numbers are allowed to mean rather than as a finding of our own: scoring extraction against that model measures how well we recovered a plot, not how close we are to the ground. We have measured nothing about the dependency ourselves, and we do not claim to have.
There is a second reason a bigger number is not a better one here. Across five sets we have exported 332 stationed cross-sections from 434 source pages. There is no denominator. Nowhere in this repo is there a count of how many sections exist on those sheets, so 332 is sections exported, not coverage and not recall. "We reached 76 more" is a larger figure with nothing under it. Any percentage we printed next to it would be invented, and the reader would have no way to tell.
Coverage is cheap to report and impossible to audit from outside. Error is expensive to measure and no one is required to report it. A shop that takes both of these fixes shows you more sections at the same stated price, and the difference shows up later as rework someone else pays for.
What it looks like when the number goes the other way
The same discipline has to apply when a measurement flatters us, or it is not a discipline.
We froze 60 elevation points before the reader that would be scored against them was written. A person read them off the sheets by hand, blind to the software's answer; the hand-check routine for reading a digitized cross-section against its own sheet is written up separately. Fifty-nine were scored. Before we corrected the measuring tool, the software came in at median 0.054 ft, p95 0.892 ft, max 3.293 ft against those hand readings. After the correction, on the same bytes, p95 0.126 ft and 58 of 59 within half a foot.
Three things about that second number, all of which have to travel with it. It is an instrument correction, not an engine improvement — the output did not change; our measuring tool had been misreading closed pavement boxes and scoring the software as wrong. The reader recorded to the nearest 0.1 ft, so roughly 0.05 of the 0.126 ft is that rounding rather than our error. And nothing was double-read, so repeatability was never tested at all. The canonical file for that check still reads FAIL, and it is deliberately kept out of the v1 gate roster.
Sixty points on one set, read by one person once, is not an accuracy claim. It is the best evidence we currently have, and it is thin.
Where this leaves the work
Every finished run we have produced — 26 of 26 — carries no surface: zero points, zero faces, marked not exportable, and marked refused. Each names the same three reasons for it: we cannot confirm which way the section is facing, the alignment is absent, and the curvature of that alignment is unknown. That is a named "no" rather than a plausible model, and it is why we would rather emit a refusal than a guess.
It is also the reason we do not claim cross-sheet agreement at all: the registered bar is 12 scoreable station comparisons, the honest count is 0, and the most favourable ground-line ruling reaches only 7 — which is why we hold controls this tightly in the first place.
The mechanics are written up separately, in how a plan set becomes machine-readable data.
The band window is still open. It is not a good change yet — it buys a larger count with a residual we cannot price. When we can tell you what the 76 sections are worth against something other than another drawing, we will turn it on and say so.
Mathyra is in private development. Figures quoted here are measurements from our own engineering runs, with their limits stated; nothing above claims an accuracy we have not shown.