Method · September 7, 2026 · 9 min
How to check a digitized cross-section against the sheet it came from
A five-step protocol for verifying digitized cross-sections against the plan set: freeze the points, read them blind by hand, score residuals, run a sign test.
Somebody hands you a surface built by digitizing cross-sections off a scanned plan set. Breaklines, a dirt model, something a machine could run. Before you let it near a blade, you want to know whether it matches the sheets it came from — and the answer to "how do I verify this" is not a report the software prints about itself.
Here is the protocol we run. It takes one person, a scale, and about ninety minutes — our reader logged an hour on 59 points plus twenty-five minutes reviewing them. The most useful thing about it is that our first run of it failed, and the failure found a real defect. That is what a check is for.
The check, in five steps
1. Freeze the points before the reader exists
Pick the points you are going to check — station, offset, and which line on the sheet you mean — and write that list down before anyone writes or runs the code that will produce the answers. We froze 60 points this way. If you pick your check points after you have seen the output, you will pick the easy ones. You will not mean to. You will do it anyway.
Freezing the list also forces you to say out loud which line you mean. On a section with several lines through the same offset, "the ground" is not a specification. Naming the point in advance makes the ambiguity visible before it becomes a dispute about the score.
2. Have a person read them off the sheet, blind
A human reads every frozen point off the printed section by hand — scale on the paper, station by station — without seeing what the software produced. Not a second script. A person, blind.
This is the step people skip, and it is the only step that puts an independent number next to the machine's number. Everything else in the protocol is arithmetic on the gap between the two.
3. Write down the reader's own precision, and read some points twice
Our reader recorded values to 0.1 ft. That is a real limit on what any residual can mean: roughly 0.05 ft of any miss we report is the reader's rounding, not the software's error.
We also ran zero double-reads. Nobody read the same point twice, so we never tested whether the same person reading the same point on the same sheet gets the same number an hour later. That is a hole in our own protocol and we are naming it rather than letting a clean-looking score cover it. If you run this yourself, double-read a fifth of your points — that is what our own pre-registration called for, 12 of 60, and what we then failed to collect.
4. Score residuals, not an average
Report the median, the 95th percentile, the maximum, and the share of points inside your tolerance band. Do not report a mean and do not report a correlation. An average hides the one point that is three feet out, and the one point that is three feet out is the one that busts grade and pays for the rework.
Say the sample size and the sample in the same breath as the number. "p95 0.126 ft against 59 hand readings on one set" is a claim you can argue with. "Accurate to a tenth" is not.
5. Run a sign test
Then look at which way the misses point. If the software is simply noisy, misses land above and below the hand reading in roughly equal numbers. If they nearly all land the same way, that is not noise — it is a bias, and a bias has a single cause you can go and find.
That is exactly how 0.44 ft of elevation bias turned out to be a font baseline: the misses were one-sided, the sign test put the odds of that happening by chance at about one in ten trillion — 56 of 59 residuals landed low, two-sided p = 1.19e-13 — and the cause was that a label's box centre and the number's baseline are not the same place on the sheet.
Our first blind reading failed, and that is the useful part
We froze 60 points as the target and the reader came back with a complete 59-row sheet — one short, with no exclusion codes used and no point thrown out. Say why a unit left your population, even when the reason is that it never arrived. A person read those points off the sheets by hand, blind to what our software said. The saved result file for that check reads FAIL, and we deliberately kept it off the list of checks we would stand behind for a first version. We are not going to tell you that check passes, because it does not.
Here is the scoring immediately before we corrected the measuring tool, alongside the same points re-scored after:
| Measure | Before the instrument correction | After the instrument correction |
|---|---|---|
| Median residual | 0.054 ft | 0.038 ft |
| 95th percentile | 0.892 ft | 0.126 ft |
| Maximum | 3.293 ft | 2.101 ft |
| Within 0.50 ft | 0.7627 | 0.9831 (58 of 59) |
Both columns are 59 hand readings on one set.
The check had already earned its keep before that column existed. In the very first scoring the worst misses — points off by seven to nine feet, against the 3.293 ft worst case shown here — were not the model being wrong about dirt. They were gridlines. The scoring was picking up the sheet's own ruled lines and treating them as ground. On a related set, gridlines accounted for 47% of what one export contained: a vertical line is not a surface, and while the export was full of them, the cap on how much it would write out was eating real terrain to make room.
There was a second, smaller version of the same problem. On 3 of 180 windows, the routine that hunts up the sheet to find the vertical scale ran off the top of its search range and admitted the sheet's template furniture — the printed frame, not the drawing. Two of the frozen points moved when that was fixed: one by -3.29 ft, one by -0.11 ft.
The second column is a correction to the measuring tool, not an improvement to the engine
The second column was produced by re-scoring the same bytes. The digitized sections did not change. Not one coordinate moved. What changed is that the thing doing the scoring had been misreading closed pavement boxes on the sheet — the measuring tool was wrong, so the measurement was wrong. We fixed the tool and measured again.
So p95 0.126 ft, against 59 hand readings on one set, is an honest number about one narrow thing — how closely the topmost drawn line on those sections matches the printed sheet. It says nothing about whether that line is the right line, about georeferenced position, or about a surface, and it is emphatically not evidence that the extraction got better. If you run a check on your own data and the score jumps, your first question should be whether the engine improved or the ruler did. Ours was the ruler. And subtract the reader's rounding before you get excited: about 0.05 ft of that 0.126 is the reader recording to a tenth.
One more disclosure that belongs with these numbers. The exports sitting on disk were not produced by the version of our engine that exists today. A score is attached to a specific build and a specific set of bytes, or it is decoration.
What this check cannot tell you
A hand reading off the sheet retires exactly one question: does the digitized section match the printed section. It says nothing about whether the printed section matches the ground, and nothing about whether it matches the designer's model.
The obvious next move is to score against the designer's own file, and on one corridor we did. Comparing our grade line to the designer's grade line gave a median of 0.325 ft and p95 0.833 ft across 86 stationed sections on that single corridor. Getting there required fixing how far sideways the software looked when reading the vertical scale; before that fix the same comparison sat at p95 3.619 ft. One corridor, n=1 — treat it as a worked example, not a rate.
But be clear about what that comparison buys. We have measured nothing here, so take this as adopted, not demonstrated: the agencies' own workflow documents describe cross-section sheets as cut from the same corridor model the bid-package file is exported from — FHWA HIF-17-031 is the clearest statement of it. If that is right, scoring one against the other measures how faithfully a drawing was plotted and recovered, not whether either one is right about the dirt, and we treat it as a binding limit on what our own numbers are allowed to mean. That distinction is the whole reason the disclaimer our own report carries reads, verbatim: "Not one coordinate in this deliverable has been compared against a surveyed value, a designer's model, or a hand-read elevation off the sheet."
Which is also why the sheet is the right first check and the wrong last one. It is available, and it is the sealed document: SUDAS Standard Specifications section 1040 says electronic support files "are for information only," and that "Should there be a discrepancy between an electronic support file and a contract document, the contract documents shall govern", Caltrans ranks "supplemental project information" last of six contract parts in Standard Specifications 5-1.02 (2025 Edition) and describes electronic design files in the subsection of that name — though whether a given project’s model carries that status is set by its special provisions, not by the Standard Specifications. Read those and your own contract; we are describing what agencies publish, not advising you. The sheet just cannot tell you about ground truth, and no amount of arithmetic will make it.
Where this stands
Mathyra is in private development. We have never shipped a surface. Every one of the 26 exports we have produced carries no surface at all — zero points, zero faces — and each names the same three reasons it refused: the left-right handedness of the section frame is unconfirmed, the alignment is absent, and the alignment's curvature is unknown — reasons that still stand because we held back a fix rather than trade one refusal for another. We would rather emit a named refusal than a number we cannot defend.
The blind reading is one check among several, and the one we point at when someone asks how you would ever know. Run it on whatever tool you are evaluating. When we have a surface worth handing anyone, run it on ours. Freeze the points first. Read them blind. Score the residuals and look at the signs. And when it fails — it will, the first time — check your ruler before you blame the engine.
If you want a case where the sheets looked like they disagreed and the fault was ours, we compared one station across the plan, profile, section and control sheets and found seven comparisons that would not reconcile — then traced the gap to our own reader pairing one profile panel's elevation ladder with another panel's ruler. We still do not claim cross-sheet agreement; our own report records it as not claimed.
Mathyra is in private development. Figures quoted here are measurements from our own engineering runs, with their limits stated; nothing above claims an accuracy we have not shown.