There are no results on this page.
The method was fixed before a single number existed, so the method could not be chosen after seeing the numbers. That is the reason this page is worth reading while it is still empty: everything on it was decided at a point where nobody yet knew which definition would flatter the engine.
What we publish is measured. What isn't measured is labeled accordingly. Three metric categories are defined below, and they are never combined into one accuracy figure. When packages have been adjudicated, each category publishes on its own terms, per rule and per rule family, alongside a section naming every case the engine missed and why.
A benchmark settled after the first run is not a benchmark.
It is a selection. The thresholds move. The awkward cases turn out to have been out of scope all along. The metric that looked worst turns out to have been the wrong metric. From outside the room, an honest result and that process are indistinguishable, and the only defence against it is to write the method down before any number exists and be held to it.
So the definitions, the adjudication protocol, the validation rules and the stated limitations of the measurement are all fixed and public while the tables are still empty. Until those tables carry rows, nothing here is a measured claim about accuracy, and this page says so in its own headline rather than in a footnote.
Published in full as BENCHMARK.md in the engine repository. Its status line, dated 10 August 2026, records the same thing: the measuring instrument exists and is tested, the dataset does not exist, and no result may be quoted.
Three metric categories. They are never combined into one accuracy number.
One blended figure would be easier to quote and impossible to defend, because the three ask different questions of different evidence and disagree with each other by design. The clearest case: a finding that causes a pre-bid RFI, which closes the gap, so no change order is ever written, cannot be scored against issued change orders. Measured that way, the engine's best possible outcome scores as a false positive.
That is not hypothetical. The scoring code that exists today counts a flag with no matching change order as a false positive, which merges "the finding was wrong" with "the finding was right and somebody fixed it before it cost anything." Separating the categories is what stops one number from being computed across both.
Finding precision
Is the finding true about the documents?
Every finding in a reviewed package is adjudicated by a person into exactly one of three verdicts, written down and signed.
- Valid
- The finding describes the documents correctly. The citation is real, the sentence says what the finding says it says, and the conflict or the gap is there. Whether anyone acts on it is a separate question, asked in category 02.
- Not an issue
- The rule fired and the documents are fine. This is the false positive, counted as one, with the rule that produced it named.
- Unclear
- The reviewer could not settle it from the documents in front of them. It stays Unclear. It is never pushed into one of the other two to make the table tidier and it is never dropped, because both of those quietly move precision.
Reported: precision, meaning Valid over adjudicated; the false-positive rate, meaning Not an issue over adjudicated; findings per project; and the review effort the set consumed, which is category 04 below. Broken out per rule and per rule family, with Unclear published as its own count in its own column and never distributed across the other two.
What it does not measure: whether a Valid finding was worth anyone's time. A rule can be right about the documents and useless to an estimator, and this category cannot tell the difference.
Instrument: protocol fixed, no scoring code written, no package adjudicated. The adjudication record described further down is category 03's, not this one's.
Commercial usefulness
Did a valid finding change what anyone did?
For every finding adjudicated Valid, two things are recorded. First, the action it caused: a pre-bid RFI, a line carried in qualifications and assumptions, a scope or price decision made differently, a clarification requested from the design team, or no action at all. No action is a legitimate and frequent answer, and it is counted rather than discarded.
Second, the counterfactual: would the estimator have caught this anyway?
Why the counterfactual is the hard part
It asks about a review that did not happen. No document settles it, no public record checks it, and there is no way to run the same package twice through the same person. It is also the only question the buyer is actually asking. Nobody wants to know whether a finding is true. They want to know whether they would have found it themselves, on a Thursday, with the bid due Friday.
So it is the weakest instrument on this page and it is the one that matters most. That is why its weaknesses are handled in the protocol rather than in a footnote printed under the result.
Why a five-point scale and not yes or no
- A binary destroys the uncertainty, and the uncertainty is the data. Forcing a reviewer who genuinely cannot say to answer yes or no does not produce knowledge, it produces a coin flip recorded as a judgement.
- Hindsight runs one direction. Once you have read the finding and opened the cited page, the gap looks obvious. A binary gives the reviewer no way to say "obvious now, and I would have walked past it at four in the afternoon the day before bid." A scale gives them that sentence.
- The distribution is the result, not the mean. The five counts are published. Averaging a subjective scale into a single number is exactly the blending this page refuses one level up.
- 1Certainly would have caught it. This is part of the check I run every time.
- 2Probably would have caught it.
- 3Genuinely cannot say.
- 4Probably would have missed it under bid deadline.
- 5Certainly would have missed it.
Recorded per finding by the estimator who reviewed it, before they are shown what any other reviewer answered, and attributed by name, because a self-reported counterfactual with nobody's name on it is not evidence. Where two reviewers see the same finding, both answers are published and the disagreement is reported rather than averaged away.
Instrument: protocol fixed, no scoring code written, no package adjudicated.
Realized outcome recall
On finished projects, did the findings correspond to what actually happened?
Run on completed historical projects only. Never on a live bid, because on a live bid the outcome has not happened yet and the engine's own output is one of the things that changes it. A package qualifies when it reached construction and its project record is closed and retrievable.
Findings are compared against five kinds of outcome, not one:
- Change orders issued. The public, itemized record from board minutes, cited to a retrievable URL and a meeting date.
- RFIs filed. The outcome the tool exists to produce. A finding that becomes a pre-bid RFI and closes the gap prevents the change order that would otherwise have proved the finding right.
- Field issues. Problems that surfaced during installation and were resolved without a change order.
- Absorbed scope. Work the contractor ate rather than argue about. This leaves the thinnest record of the five, and it is the reason precision measured against change orders alone is a lower bound.
- Documented disputes. Claims, back-charges and disagreements recorded in the project file.
Reported: true positives, false positives and false negatives; precision, recall and F1; and lead time as a distribution carrying minimum, median, mean and maximum in days plus the per-match detail. Median is the headline lead time rather than mean, because one change order issued three years after award drags a mean past anything a bidder would recognise. Broken out per rule id and per rule family, with unmatched flag ids and unmatched change order ids listed by name rather than summarised into a count.
Instrument: the change-order half is scored today by bench/risk_score.py, which is checked in, typed and tested. No package has been adjudicated, so it has never produced a result. RFIs, field issues, absorbed scope and documented disputes are specified here and are not in the adjudication record format yet, which is labeled rather than left to be discovered in an empty column.
The rule, stated once so it governs every future version of this page: these three never collapse into a single accuracy figure, a single score or a single percentage in anything this project publishes. A finding can be true and useless. A finding can be useful and leave no trace in the public record. A finding that prevented a change order from ever existing cannot be scored against issued change orders. Three questions, three answers, published separately.
What the output costs the person who has to read it.
A register nobody has time to read is worth nothing however precise it is. The cost is one open-the-page-and-read-it per flagged line, paid by the estimator, in the week they have the least time. It is measured in minutes, and the measurement is specified here with the rest of the method rather than added later if the result happens to look good.
| Measurement | Definition | What it answers |
|---|---|---|
| Total review minutes | Wall-clock time one reviewer spent adjudicating one package's register, recorded per finding and summed. | What the whole exercise costs on one package. |
| Minutes per finding | Total review minutes divided by findings adjudicated, valid or not. | What one line of output costs to dispose of, including the lines that turn out to be nothing. |
| Minutes per valid finding | Total review minutes divided by findings adjudicated Valid. | What a true finding costs once the false positives are paid for. |
| Minutes per actionable finding | Total review minutes divided by findings adjudicated Valid that also caused an action under category 02. | The one that decides whether the tool is worth opening. |
The eventual metric, stated plainly: how much estimator attention Blueprint consumes per useful issue it surfaces.
No number for any of the four appears here, and none will appear until it has been timed on real packages by real reviewers. There is no modelled figure, no estimate and no illustrative example, because an illustrative number in this position is the number that gets quoted back. The only run anyone can point to today is the repository's own hand-authored twelve-opening fixture, and quoting a synthetic fixture as an expected workload would be dishonest.
The denominators matter as much as the totals. A ratio with a zero denominator renders n/a and never 0.0, here as everywhere else in the harness: a package that produced no actionable findings did not cost zero minutes per actionable finding, it produced no such measurement at all.
Stated here in full, once, and pointed to from everywhere else.
This engine has not been measured against real outcomes. There is no benchmark dataset, no precision figure, no recall figure and no accuracy claim anywhere on this site. Expect both false positives and false negatives. How often either occurs is unknown, because nobody has counted yet.
Nothing in the output, the rule packs, the plan viewer or the risk register may be read as evidence that the engine is correct. Those artifacts make the output auditable. Auditable and correct are different properties and only one of them has been demonstrated. Nobody should use this as the only review of a bid package, and the design does not ask to be: it performs one cross-check that gets skipped under deadline and hands back a list of questions for a person to answer.
This caveat appears in full in exactly two places on this site, here and on the limitations page. It is not repeated on every screen, because a warning printed in the margin of every page stops being information and becomes decoration, and a reader learns to skip it. Stated once, in the two places a reader goes looking for it, it keeps its meaning.
Recorded in the engine repository as amendment A-003. The constraint is permanent and is not waived by any later phase.
Three layers of ground truth. Two of them have scoring code today.
The third has none, and its absence is labeled rather than implied by an empty table. Every metric below is one the scoring code actually computes; nothing here is a metric somebody would like to have.
| Layer | Ground truth | Metrics | Scoring code |
|---|---|---|---|
| Risk truth | Change orders actually issued, read from public board minutes | Precision, recall, F1, lead time | bench/risk_score.py |
| Extraction truth | Hand-annotated drawing sheets | IoU for walls, IoU for rooms, precision and recall for openings | bench/extraction_score.py |
| Design truth | The equipment schedule published in the bid package | Device count error by category, hardware type match rate, power supply and battery quantity accuracy, false-omission rate | Planned |
Risk truth, in full
The scorer emits true positives, false positives and false negatives; precision, recall and F1; and a lead-time distribution over matches carrying the minimum, median, mean and maximum in days plus the per-match detail. False negatives are split into attributed and unattributed. Breakdowns are produced per rule id and per rule family, and the unmatched flag ids and unmatched change order ids are listed by name rather than summarised into a count.
This is the change-order half of category 03, and it is the only instrument of the four on this page that runs today.
Extraction truth, in full
Wall IoU is aggregate: the union of every annotated wall polygon against the union of every predicted one, giving one intersection, one union, one number. Rooms are scored twice, aggregate and mean over rooms matched one to one, with the aggregate as the headline because a mean over matched rooms is trivially gamed by finding one room perfectly and missing twenty. Openings carry precision, recall, F1, and a mark match rate reported separately from all three.
The two matching parameters, an opening tolerance of 2.0 feet and a room IoU threshold of 0.5, are carried on every report so a published number cannot be read without them.
Design truth is Planned because plan extraction is Planned, so there is nothing to compare a published equipment schedule against. When it is built, a package's own equipment schedule is registered as truth and is never fed to the engine as an input for the same package it grades. A system that reads the answer key cannot be scored against it.
A person decides whether a flag and a change order are the same event.
Whether a pre-bid flag and a realized change order describe the same thing is a judgement about construction documents. It is made by a human reviewer, written down, and signed. The scoring script consumes that file and computes arithmetic over it. No model performs a match anywhere in this benchmark, and an architecture test walks the import graph and fails the build if a model client is reachable from the harness.
True positive
A flag adjudicated onto a change order. The adjudicator has to name both and write down why they are the same event.
False positive
A flag no change order matched. Read it against category 01 before reading it as an error: this count includes correct findings that were resolved before they cost anything.
False negative
A change order no flag matched. This is the miss, and it is the number the failures section of the document exists to explain one case at a time.
One JSON file per package, format version 1.0.0
The flags, the change orders and the matches live in one file rather than three, because a match list is only meaningful against the exact flag set and change order set the adjudicator had in front of them. Splitting them makes it possible to score last month's human decision against this morning's re-run, which silently converts a judgement into a claim nobody made.
- flags
- Every pre-bid flag the engine produced on this package, reduced to what scoring needs: flag id, rule id, the target entity and its label, the finding, the date it was flagged, and the basis for that date. Ids must be unique.
- change_orders
- Every change order the district actually issued, each with its identifier, what it did in the record's own terms, the scope it touched, the date issued, and a citation naming the public URL and meeting date it was read from. A citation that is not a retrievable URL is rejected at parse time. A change order asserted without one is folklore, and folklore in a benchmark is worse than a missing row.
- matches
- Each pairing of one flag to one change order, carrying the adjudicator's name and the date, and a note explaining the decision. The note is what makes the match reviewable: without it a later reader cannot tell whether the adjudicator matched on the substance or on the fact that both mention doors. A blank note is refused, not warned about.
- bid_opening_on
- The line between prediction and hindsight. Everything the engine may be credited with had to be knowable before this date. Every flag must be dated no later than it and every change order no earlier, which is what makes lead time non-negative by construction rather than by a check.
- engine_output_sha256
- The hash of the engine output the flags came from, so a result can be tied to the exact run that produced it and a re-run that changed the register is detectable rather than invisible.
- change_order_record_complete
- True means the minutes below were read through and the change order list is what they contain. False means the research is unfinished. An empty list under False is refused outright: a project nobody has researched and a project that genuinely issued no change orders must not produce the same recall.
- minutes_reviewed
- Every public record searched, including the ones that yielded nothing. Claiming a complete record with no change orders requires citing the minutes that were read and found empty, because "we looked and found none" is a claim about specific documents and has to name them.
A flag's date is the issue date of the latest document it cites, so a flag argued from Addendum 3 is not available before Addendum 3 exists. This record covers category 03. The record formats for categories 01 and 02 are specified in prose above and are not written yet.
Every one of these is a refusal to produce a number, not a warning.
A benchmark can be gamed quietly and in good faith. These are the specific ways this one could have been, and the specific thing that stops each.
A change order carries no monetary field
The record has no place to put a dollar figure and none may be added. The loader goes further and rejects any payload key that looks like one, with a message naming this rule rather than reporting an unknown field, because a key rejected as a typo invites the reader to add it back.
The reason is that a change order's value measures labor rates, prevailing wage, overhead, profit, and how badly that contractor wanted the job. The same scope gap on the same building in the same month is priced differently by a multiple depending on whether the contractor is carrying an idle crew or buying overtime, and none of that variance is about whether the engine found the gap. Scoring against dollars would also let one expensive change order swamp twenty correctly predicted cheap ones.
A ratio with a zero denominator renders as n/a
Never as 0.0, and never as a blank. Precision over zero flags is not zero precision, it is no precision, and a blank cell in a published table gets quoted as a zero by whoever reads it next.
"The engine was wrong about everything" and "the engine said nothing" are different results and must never print the same. F1 is undefined wherever either input is undefined, and zero only where both are defined and zero, because a run that produced flags, had change orders, and matched none of them really did score zero. That is a measurement, not a gap.
The engine output hash is recorded
Every adjudication names the sha256 of the engine output its flags came from, so a published result can always be traced back to the run that produced it. Both scoring scripts are pure functions of their inputs: they read no clock, make no network call and hold no state, so the same files produce byte-identical output on any machine.
One flag, one change order, in both directions
A flag matches at most one change order, because matching it to two would count a single correct prediction twice and inflate precision. A change order is matched by at most one flag, because matching it twice would let a family of overlapping rules turn one realized outcome into several true positives.
A miss is attributed to a rule only by a human
An unmatched change order is counted against a specific rule only where the adjudicator wrote down which rule ought to have caught it, and why. Nothing is inferred. A change order nobody attributed still counts in the overall recall denominator and is reported separately as unattributed, so per-family recall can never be flattered by quietly dropping the misses nobody assigned to a family. Both denominators are published.
An empty file is refused outright
A package with no flags and no change orders produces no report. There is nothing to measure, and reporting zeroes would make "no data" and "the engine found nothing" print identically. The breakdown is published per rule and not only per family, because a family-level number hides which rule is carrying it and which is sinking it, and the per-rule table is the one that says which rule to fix.
The method has stated limitations of its own, and they will still be true when the tables are full. Public bid packages carrying both Division 08 and Division 28 that also reached construction with a public change-order record are rare, so the first result will rest on a small number of projects, and a number computed over five change orders is a number, not a rate. Districts that publish legible itemized minutes are districts with the administrative capacity to produce coordinated bid documents in the first place, so the measurable population is not the target population. A gap the contractor absorbed, or resolved through an RFI, leaves no public record and scores as a false positive under category 03, which is the precise reason category 01 is measured separately. And one adjudicator is one opinion: independent re-adjudication by a second person, reporting the disagreement rate, is the correct next step and has not been done.
The rule families whose numbers this page will carry.
Every breakdown here is published per rule id and per rule family, so the families have to be identifiable, versioned and challengeable before any number is attached to them. Each is a JSON file in the engine repository carrying its own version, its jurisdiction and the technical authority each rule cites.
| Family | Rule ids | Technical authority | Version | Last changed | Change history |
|---|---|---|---|---|---|
| Scope boundary scope_boundary |
SB-001, SB-002 | CSI MasterFormat Division 08 and Division 28 boundary; AIA A201 General Conditions 3.2 on the contractor's review of the contract documents; public contract code pre-bid clarification practice | 0.1.0 CA-DSA |
9 Aug 2026 853ee72 |
git log packs/scope_boundary.json |
| Division seam division_seam |
SB-020 to SB-024 | CSI MasterFormat Division 08 and Division 28 boundary; AIA A201 General Conditions 1.2.1 on complementary documents, 3.2 on pre-bid review, and 3.2.2 on reporting discovered discrepancies before proceeding | 0.1.0 CA-DSA |
9 Aug 2026 853ee72 |
git log packs/division_seam.json |
| Access control ac_k12_base |
AC-001, AC-010 to AC-014, AC-020 to AC-022, AC-030, plus validators V-001 to V-006 | NFPA 101 Life Safety Code 7.2.1.5; CBC 1010.2 and 1010.2.13; UL 294 Access Control System Units; ANSI/BHMA A156.31 Electric Strikes, Grade 1; NFPA 72 10.6 Secondary Power Supply; TIA-568 horizontal cabling distance; CSI 28 13 00, 08 41 13 and 08 71 00 | 0.1.0 CA-DSA |
26 Jul 2026 78030a9 |
git log packs/ac_k12_base.json |
| Video surveillance vs_k12_base |
VS-001, VS-002 | IEC 62676-4 Video surveillance systems, operational requirements; CSI 28 23 00 Video Surveillance; district security design standard | 0.1.0 CA-DSA |
26 Jul 2026 1b1f515 |
git log packs/vs_k12_base.json |
External review: not yet commissioned
No architect, code consultant, authority having jurisdiction or independent security engineer has reviewed these rule packs. The field is here, labeled and empty, rather than absent, because a governance table with no reviewer row reads as though the question was never asked. When a reviewer is engaged, their name, their standing and the date go in this position, and the rules they disputed go in the failures section.
Ids are permanent and results cite them
A rule id is never renumbered and never reused. A superseded rule keeps its id, stays resolvable, and gains a pointer to the id that replaced it, because a published benchmark result cites rule ids and a renumbering would silently reattach an old number to new behaviour. Rule ids are the join key between this page and the register.
Every rule carries a rationale written for a practitioner and an authority citing a code section or a manufacturer document, and both print in the deliverable. Every rule is project and authority-having-jurisdiction overridable: the packs are data, not compiled behaviour, and a rule that reads a code section the way your jurisdiction does not is meant to be turned off and argued with.
Read the published rulesThe dataset is the long pole, and it is assembled by hand.
The instrument was the easy part. What the benchmark needs is adjudicated change order records from real bid packages, obtained from public sources, and nothing enters the dataset until a human has opened the source document and read it.
Public portals and board minutes
PlanetBids district portals, district purchasing and business-office pages for solicitations, addenda and bid tabulations. BoardDocs and Granicus agendas and minutes for the change-order records themselves.
Anything behind a login
Documents requiring a login or a vendor registration are not retrieved. That constraint costs coverage, and it is kept anyway, so that a third party can check every claim made here against the same documents.
Manifests only, never the documents
A manifest records the bid number, the district, the source URL, the retrieval date and the derived annotations. Source bid documents, drawings and client material are never committed to the repository.
The scarcity is real and it is not going to be solved quickly. A usable package has to carry both Division 08 and Division 28, reach construction, and have a district that published legible itemized change orders afterward. Then a person has to read the minutes, write the record, and sign the matches. Until enough of those exist, there is no result, and there will be no estimate of one here in the meantime.
The code that will produce the numbers has its own suite.
The scoring harness is not a script written the week a result is needed. It is checked in, typed, and tested against its own refusals, so that the first published number is produced by code that already exists and has already been argued with.
tests under tests/bench/
Covering both scoring scripts. They exercise the refusals as much as the arithmetic: that a monetary key is rejected by name, that an undefined ratio renders as n/a and never as a number, that a flag cannot be matched twice, and that scoring the same file twice produces byte-identical output.
mypy covers bench/ as well as src/bre
The type checker is configured over both trees, not just the product code, and it is clean. The benchmark harness is held to the same standard as the engine it scores, because a result produced by unchecked code is a result nobody can defend when it is questioned.
Both scripts run from the command line against a single adjudication or annotation file and render markdown, so the tables in the published document are produced by code rather than typed in by hand. A result cannot be entered in a form the scorer never produced.
If your packages could contribute adjudicated records, that is the fastest way this page gets numbers on it.
A pilot participant who has run a package that reached construction, with a district that published its change orders, holds exactly the material this benchmark is short of. The engine output and the public record are what get adjudicated; no pricing is recorded and none is scored.