← Field notes

Method · September 7, 2026 · 11 min

We could not find an accuracy number for plan-derived models — and ours would be wrong too

We could not find a published accuracy figure for PDF plan sets turned into 3D surfaces. Why scoring against the designer's model measures the wrong thing.

We could not find a published accuracy number for turning a PDF plan set into a 3D surface — not from a DOT, not from a software vendor, and not from us. Search for one and you get 3D-printing tolerances and apps that trace residential floor plans. The gap is real, and it does not persist because nobody thought to measure. It persists because the cheapest way to produce that number produces a wrong one — and we would produce the same wrong one if we scored ourselves the obvious way.

Here is the obvious way. Find a bid package that ships both a plan set and a design model. Digitize the cross-sections. Compare your surface against the designer's LandXML. Publish the difference and call it accuracy.

That number measures how well you recovered a plot. It does not measure how close you are to the ground.

The disclaimer we print, and the three things it says are unchecked

Every model package our engine produces carries this sentence, verbatim:

Not one coordinate in this deliverable has been compared against a surveyed value, a designer's model, or a hand-read elevation off the sheet.

Three legs. A surveyed value. A designer's model. A hand-read elevation off the sheet. A sheet-versus-LandXML comparison retires exactly one of them — the designer's model. It cannot touch the survey, and it cannot touch the hand reading. Any vendor who runs that comparison and calls the result accuracy has retired one leg of three and quietly implied all three.

Why the middle leg is the weakest one to lean on

The agencies' own documentation describes cross-section sheets as plotted from the corridor model and the bid-package LandXML as exported from that same corridor model. We have not measured that ourselves — zero measurements — and we adopted it as a standing limit on what our own numbers are allowed to mean rather than as a finding of ours. Scoring one against the other is then a round trip through the plotter: you are asking how much a drawing loses when it is printed and read back, not how much either one differs from the dirt.

Here is what those agencies say. ODOT's RCM-400 has typical sections drawn with the vertical scale exaggerated — a rendering choice made for legibility, not a measurement. TxDOT’s PS&E Preparation Manual posts cross sections and 3D models under a section headed "For Information only," with a mandated disclaimer that the data "is for non-construction purposes, only", alongside its Model Development Standards and its Model as Legal Document work. Caltrans ranks "supplemental project information" last of six contract parts in Standard Specifications 5-1.02 (2025 Edition) and describes electronic design files in the subsection of that name — though whether a given project’s model carries that status is set by its special provisions, not by the Standard Specifications. FHWA’s technical brief "Utilizing 3D Digital Data in Highway Construction" (FHWA-HIF-17-031, April 2017) makes an engineering point rather than a contractual one: "The data is often not sufficient for construction due to a variety of reasons. The most notable is that the original ground basis for the design differs to field conditions." The same brief recommends design practices "that prioritize the 3D model as the source of the contract plans," which cuts toward elevating the model, not subordinating it. WisDOT's Construction Data Packet requires the contractor to build to the plans. SUDAS says electronic support files "are for information only," and that where one disagrees with a contract document, "the contract documents shall govern".

Read those together and the picture is not consistent, which is the useful part. Most of these agencies subordinate the model — but PennDOT defines Model as the Legal Document, under which "a model(s) comprises the primary construction contract document, preeminent in importance," and Iowa’s own precedence list ranks digital contract files above the plans where they exist and the contractor uses automated machine guidance. The answer is not "the plans always govern." It is that the answer is set per agency, and sometimes per project.

The one corridor we have ever scored this way

One. Not a sample — one corridor, n=1.

It is the FHWA Federal Lands Highway project OR FLAP DOT CRGNSA 100(9), Historic Columbia River Highway. Public domain, pulled by hand from SAM.gov on 2026-09-04. Oregon here is the location, not the owner; it is not Oregon DOT.

The paired model gives 45 CoordGeom elements over 5,501.409 ft with a maximum join gap of 0.0000 ft, a finished TIN of 13,436 points and 26,388 faces, and an existing-ground TIN of 48,694 points and 92,582 faces. Against 90 stationed sections we scored 1,034 lines across nine pen families, excluded 1,212 as sheet furniture — borders, title blocks, grid and other printing that is not ground — and refused 1,229 for having too few samples to score.

ComparisonSections within 0.50 ftMedianSample
Our grade line vs the designer's grade line0.325 ft (p95 0.833 ft)86 sections
Solid pen family vs finished-grade TIN19 of 80 (23.7%)RMS 1.503 ft80 scored sections
Solid pen family vs existing-ground TIN14 of 81 (17.3%)RMS 1.688 ft81 scored sections
Hairline dashed family vs finished-grade TIN20 of 61 (32.8%)RMS 0.823 ft61 scored sections
Hairline dashed family vs existing-ground TIN55 of 86 (64.0%)RMS 0.15 ft86 scored sections

Two things there matter more than the medians.

The first is that the numbers moved when we fixed ourselves, not when the drawing changed. On the same sheets, correcting how the reader scopes the vertical scale took the grade-line comparison from p95 3.619 ft to p95 0.833 ft. A number that moves that far on a tooling change is a number about the tool.

The second is what the comparison did and did not answer. It proposed which way the offsets run: RIGHT beat LEFT 85 to 4, with the winning convention at median RMS 0.15 ft against the loser's 1.085 ft — a proposed row in the decision log, not a confirmed one, which is why every delivery still refuses for unconfirmed frame handedness. Every semantic class abstained outright. There are zero confirmed rows in the decision log: the paired model never told us which drawn line is finished grade and which is existing ground. It proposed a handedness and declined to name a surface.

The blind reading is the only check that has ever caught a bias in the elevations

Sixty points were frozen before the reader that would score them was written. A person read them off the sheets by hand, blind to what the software said. Fifty-nine were scored.

Before we corrected the measuring instrument: median 0.054 ft, p95 0.892 ft, max 3.293 ft, 76.27% within 0.50 ft, 72.88% within 0.25 ft. After the correction, on the same bytes: p95 0.126 ft, 58 of 59 points, 98.31% within 0.50 ft.

That was an instrument correction, not an engine improvement. Nothing about the model package changed. The measuring tool had been misreading closed pavement boxes. Separately, the 7-to-9-ft errors that check had been reporting turned out to be gridlines the check itself was scoring as ground.

So p95 0.126 ft is not an accuracy figure and we will not present it as one. It is agreement with one person's hand readings on one set. The reader recorded values to 0.1 ft, so roughly 0.05 ft of that 0.126 is the reader's own rounding. No point was double-read — the count is zero — so repeatability was never tested. The canonical artifact for that check reads FAIL, and it sits deliberately outside our v1 gate roster.

It is still the most valuable check we have, because it is the only one that has ever caught a defect in what the sheet says rather than in how our reader measures — the LandXML comparison caught how our reader was scoping the vertical scale, but only the hand reading could catch a bias in the elevations themselves. It is what surfaced the half-foot that was not there, a systematic elevation bias traced to a font baseline. Comparing against the designer's model has never caught anything of that kind, and on the agencies' own description of the workflow we do not expect it to: both sides of that comparison are described as descending from the same corridor.

The sheet does not tell you which line is the dirt

A pen family is a width, a colour and a dash pattern — what the drafter's pen was set to, not what the pen drew. If a plan-to-model process is going to build a surface, something has to decide which of those lines is finished grade.

On the Oregon cross-sections, sections carry between one and seven distinct line styles: 19 sections have one, 22 have two, 13 have three, 12 have four, 8 have five, 9 have six, 3 have seven. Only 19 of 86 are unambiguous on style alone. On FL-003 it is worse — every one of its 111 sections carries between 8 and 15 distinct styles, most commonly 10, 11 or 12. Zero of 111 are unambiguous.

The legend does not rescue this. On this corpus, legends do not map line styles to surfaces, and the layer names that would disambiguate them live in the PDF's layer metadata rather than printed on the sheet. This is the part of the problem that no accuracy number addresses, because getting it wrong does not produce a slightly-off surface. It produces a confident surface of the wrong thing.

Where the defect is ours

The routine that stitches drawn strokes into a continuous line joins them by endpoint proximity and never checks that the result is still a single elevation per offset. Thirty-seven percent of one family's lines come out multi-valued.

The worst case we have measured is Oregon page 21, station 40575. A ground line traces out and back over the same offsets, then drops 64.04 ft straight down at offset 39.83 where it terminates on the plot frame — all inside one exported line. Across the corpus that produces 305 refusals for the same reason: two elevations claimed at one offset. We tried splitting every line at its reversals; the count went from 305 to 305. The collisions are between the pieces, not inside them.

We also lose data we have not yet recovered, and we say so. On one 56-page cross-section volume, the routine that locates each section's horizontal scale dropped 100 of 211 of them. A candidate fix reaches 162 sections, but the 76 newly reached sections score a median 1.106 ft against the 84 carried-over sections' 0.323 ft. That is coverage bought with quality, so the fix that reached 76 more sections was deferred rather than traded and a test pins the deferral. We made the loss visible. We have not fixed it.

What to ask a vendor for instead

If somebody quotes you a tolerance for a plan-derived model, these are the questions that separate a measurement from a marketing number.

1. Against what? Survey, designer's model, or hand reading. If it is the designer's model, ask whether the sheets and the model came from the same corridor. They usually did. 2. What is the denominator, and what left it? A count of sections produced is not coverage. We export 332 stationed cross-sections across five plan sets (what 434 pages of plan sets actually yielded, with the denominator stated), and we are explicit that there is no denominator of sections that exist — 332 is what came out, not what was there. It is also a count from deliveries the current engine did not produce, and we say that too. In our own status reporting, 56 of 104 cross-sheet comparison units are absent across four sets with no recorded reason, and we publish that as inconclusive rather than netting it out. 3. How many corridors, not how many points? One corridor is n=1 no matter how many points came out of it. Ours is n=1 too. 4. Did a person read any of it blind? Hand-read checks catch the errors that self-consistent software cannot see in itself — how to check a digitized cross-section against the sheet it came from is the procedure we run. Ask for the count, and ask whether any point was read twice. 5. Which line did you call finished grade, and how? Then ask what happens on a section carrying twelve distinct line styles and a legend that does not name any of them. 6. What does the tool do when it cannot tell? Ours writes a named reason instead of a number — that is the whole design, and why we refuse to guess is the longer version of the argument. A tool that always returns a surface is not more capable; it is less honest about the same uncertainty.

The one question we cannot yet answer for ourselves is the first one, at any scale worth quoting. We have acquired six paired plan sets and design models from FHWA Federal Lands Highway via SAM.gov — 663 plan pages plus 81 cross-section pages, 188 LandXML surfaces, 69 alignments and 1,905,822 points, though one set contributes 81% of those points, only three of six carry both an existing-ground and a finished-grade surface, and only two of six carry an EPSG code in the header. Nothing has been extracted from it yet. It is an acquisition, not a result — and even when it is a result, it retires one leg of three.

Where this leaves us

Mathyra has never shipped a surface. All 26 model packages the engine has produced refuse to export one, every one of them for the same three stated reasons: the frame handedness is unconfirmed, the alignment is absent, and the alignment curvature is unknown. In our own status reporting, cross-sheet agreement is not claimed — the registered bar was 12 scoreable station comparisons and the honest count is 0, with the most favourable ground-line ruling reaching 7. External accuracy is unmeasured, and the report says so in those words.

That is a worse-sounding position than a tolerance figure, and a more useful one. The same station on four sheets walks one station through plan, profile, section and control, which is where this problem is easiest to see. Mathyra is in private development: no customers, no pilots, and no accuracy number — including the one we could have published this week.

Mathyra is in private development. Figures quoted here are measurements from our own engineering runs, with their limits stated; nothing above claims an accuracy we have not shown.