Limitations
What we publish is measured. What isn't measured is labeled accordingly. This page is the complete list: what Blueprint does not detect, what it gets wrong, and what its numbers are not.
There is no accuracy figure on this page because there is no accuracy figure. What is here instead is the exact shape of the gap, the measurement that would close it, and a capability table stating which parts of the design run today and which are still on paper. It is maintained the way a defect list is maintained, and it is checked against the code rather than against the roadmap.
Stated here in full, once, and pointed to from everywhere else.
This engine has not been measured against real outcomes. There is no benchmark dataset, no precision figure, no recall figure and no accuracy claim anywhere on this site. Expect both false positives and false negatives. How often either occurs is unknown, because nobody has counted yet.
Nothing in the output, the rule packs, the plan viewer or the risk register may be read as evidence that the engine is correct. Those artifacts make the output auditable. Auditable and correct are different properties and only one of them has been demonstrated. Nobody should use this as the only review of a bid package, and the design does not ask to be: it performs one cross-check that gets skipped under deadline and hands back a list of questions for a person to answer.
This caveat appears in full in exactly two places on this site, here and on the benchmark page. It is not repeated on every screen, because a warning printed in the margin of every page stops being information and becomes decoration, and a reader learns to skip it. Stated once, where a reader goes looking for it, it keeps its meaning.
Recorded in the engine repository as amendment A-003. The constraint is permanent and is not waived by any later phase.
Everything below follows from the same discipline. What is measured is published with its method. What is not measured is labeled, in one vocabulary, with the same four build states used everywhere on this site. The rest of this page is the complete list of what the engine does not detect.
The measurement method, published before any resultExpect both false positives and false negatives.
That sentence is the honest summary of what this tool is, and it is at the top rather than in a footnote because it is the thing that decides whether the tool is any use to you.
A false positive is a flag on a line where nothing is actually wrong. The rule fired, the citation is real, the sentence it quoted says what it says, and there is still no problem. You open the page, read the line, and close it. That costs you a minute and it costs the tool credibility, which is the more expensive of the two.
A false negative is the gap the engine walked straight past. It is worse and it is quieter. A short register looks like a clean package, and the two are not the same thing. There is a known example already recorded rather than hidden: the scope detector matches marker phrases such as NIC, by others, furnished by, installed by, OFCI and CFCI, and on the first real project manual it was run against it missed a page that assigned owner-furnished scope in ordinary English with none of those words in it. That page was the residual catch-all, which is the most expensive form of the thing it was built to catch.
That gap was left open on purpose. Widening the pattern to catch phrases like "responsible for" would fire on a large fraction of any specification's body text, and a rule that fires too broadly is worse than an absent rule because it trains you to ignore the output. The gap is held open, recorded, and guarded by a test that fails if anyone closes it quietly.
Sources of false positives
- Rule scope. A rule written for one jurisdiction's reading of a code section fires in a jurisdiction that reads it differently. Rules cite their authority and stay project-overridable for this reason.
- Genuinely resolved language. A scope marker that a later sentence correctly settles can still be located and flagged.
- Match error. Two lines that look like the same opening and are not.
- Uncalibrated ranking. Exposure bands are placeholders, so an item can sort higher than it deserves.
Sources of false negatives
- Vocabulary the detector does not hold. The recorded case above. Scope stated in plain English is not matched.
- Rule coverage. Several rule families are allocated and ship no rule. The table below names each one.
- Documents it never reads. Drawings are not read at all. Addenda are not reconciled.
- Evidence it cannot see. Half the Division 08 to Division 28 seam analysis reads structured claims, and nothing populates those on any package the tool can currently produce.
This framing is borrowed openly from how static analysis tools document themselves. A security scanner has the same structural problem: it fires on code that is fine and stays silent on code that is not, and the only usable posture is to say so at the top and let the reader decide what the review is worth. Nothing about construction documents makes that problem easier.
It has not been benchmarked
There is one measurement that would settle whether this engine is worth anything, and it does not exist.
That measurement is precision and recall of pre-bid flags against the change orders a project actually issued, plus lead time: how far ahead of the event the flag was raised. Precision is what fraction of the flags mattered. Recall is what fraction of the things that mattered got flagged. Lead time is whether knowing would have helped.
Nothing else substitutes for it. Not a passing test suite, not a deterministic replay, not a well-cited register. Those make the output auditable. They say nothing about whether the things it flags are the things that become change orders.
The method intended to produce it
Published here before there are results, because a methodology published after the numbers is a methodology chosen to fit them.
- Source packages where the outcome is public. Bid documents from district procurement portals and purchasing pages, paired with the change-order record from that project's own board minutes. Both halves have to be publicly retrievable by a third party, from a citation, without going through this project.
- Store manifests, never documents. The dataset holds the bid number, the district, the source link, the retrieval date and the derived annotations. Source bid documents and drawings are never committed. Anyone checking the work retrieves the originals themselves.
- Require an additive change order traceable to a document defect. This correction came out of the sourcing search and it matters more than any lead it produced. Most publicly documented school security change orders are deducts: an unused allowance returned at closeout, equipment specified and not installed, a time extension for a supply delay. A deduct has no pre-bid document defect underneath it, so recall measured against one is undefined. What scores the engine is a change order that added scope for a stated reason: work assigned to nobody, a count wrong in the schedule, a division seam neither section covered, a pathway or power requirement nobody priced.
- Confirm the project is actually in scope before accepting it. In school board agendas, "safety and security" usually means Division 27: intercom, paging, bells, clocks, classroom audio. It does not usually mean Division 28. Scanning agenda titles produces a steady stream of these, and one was accepted as a leading candidate here before the specification was opened and found to contain zero occurrences of Division 28. No candidate is accepted until a Division 28 section number or access control hardware is confirmed in the specification itself. That check costs one download.
- Score counts and outcomes, never awarded dollars. An award price carries labor rates, prevailing wage, overhead, profit, and how badly that contractor wanted the job. It measures a bid strategy, not a document defect.
- Publish the failures with the results. Per-category breakdown, and a section listing every case the engine missed and why. The failures section is the most credible thing the document will contain.
Where that stands today
No package satisfying all three sourcing criteria has been found in public sources. The failure has a clean shape: both halves exist, and not in the same project. The districts that publish full bid documents online have tended to be running Division 27 work. The districts that ran Division 28 projects hold their bid documents in a business office for in-person review.
So the status is not "the benchmark is in progress." The instrument exists: two scoring scripts are checked in, typed and covered by 89 tests, and they have never produced a result because no package has been adjudicated. The dataset is the missing half, and the results tables are empty. When they carry rows, they go on this site with the misses in them.
The method has since been fixed in more detail than this section carries, in three metric categories that are never blended into a single accuracy number, plus the measurement of what the review costs the estimator who reads it.
A methodology published before the results is a commitment. Published after, it is a defense.
The three metric categories, in fullWhat runs today, and what does not.
Several things a reader would reasonably assume are in the box are not. Below is the capability list, in the four states used on every page of this site, checked against the code rather than read off module names. One vocabulary, four values, and no fifth: "not built", "not wired", "not validated" and "not ready" were four ways of saying the same thing in four different tones, and a page carrying all four reads as a product apologising for itself.
Available Runs today on a real package.
Pilot Runs, and is being validated with design partners.
In Development Being built, not reachable yet.
Planned Specified, id block reserved, no code.
| Capability | What that means in practice | State |
|---|---|---|
| Deterministic rule evaluation, validators, risk register, report | Rules resolve in priority order under conflict groups, validators return computed values rather than a boolean, the register ranks deterministically, and the report and CLI run end to end. Golden tests compare output byte for byte. No model call anywhere in it. | Available |
| PDF text and CSI segmentation | Reads a package of text-layer PDFs, segments by CSI MasterFormat section numbering, and carries document, page and coordinate provenance on every element. No model call, no network, no clock. | Available |
| Scope marker detection | Regex over marker phrases, not a model. The recall gap described at the top of this page is this component. It has been run against one real project manual and the result was recorded rather than assumed. | Available |
| Specification and project-manual cross-check, Division 08 to Division 28 seam | Reads two kinds of evidence: located scope statements and structured claims. Only the first is populated on a package the tool can currently produce, so half of the seam analysis has nothing to read. | Pilot |
| Scope-ownership statements, SB-001 and SB-002 | One located scope statement at a time, joined to nothing else. A device the documents assign to another party, and a device the documents assign to nobody, are separate findings with separate severities. | Pilot |
| Seam rules SB-020 to SB-024 | For one opening and one aspect at a time: locking device, power, rough-in. Both divisions, neither division, conflicting requirements, and the case where the documents do not settle it, which is a distinct rule so it is never reported as the case where nobody is assigned. | Pilot |
| Access-control opening rules, AC-001, AC-010 to AC-014, AC-020 to AC-022, AC-030, with validators V-001 to V-006 | Base kit per controlled opening, egress resolution, storefront withholding, controller allocation, supply sizing and battery standby. Every rule ships a positive test proving it fires and a near-miss negative test proving it does not. | Pilot |
| Video-surveillance rules, VS-001 and VS-002 | Coverage classing and coverage-gap detection. Labeled Experimental: the pixels-per-foot thresholds and the coverage fraction are configurable planning figures, and video surveillance is not what this engine is aimed at. It is not a homepage promise and it should not be read as one. | In Development |
| Structured claim extraction | The one sanctioned model call for reading specification text is written and unit tested. No command reaches it, and the ingest path never populates the claims field. It is code that has never run on a real package. | In Development |
| Dual-vendor extraction agreement | The mechanism that reads the same content with two vendors and emits no value where they disagree is written and unit tested. It is reached by no command, and the rule id that would carry a disagreement into the register is defined by no rule pack. | In Development |
| Hosted web application, login, chat copilot | Grounded chat over one loaded package exists in the repository and its retrieval is code, not a model choosing what to read next. Nothing is deployed. There is no application you can sign into and no upload feature. | In Development |
| Schedule-to-specification join, SB-010 to SB-019 | The block for "a device appears in a drawing schedule and every specification reference assigns it elsewhere, or nowhere" is reserved and empty. The schedule-row record is modelled, nothing populates it, and no pack defines those ids. The project contract calls this the highest-value check in its phase, and it does not run. | Planned |
| Addendum reconciliation, DC-020 to DC-029 | There is no addendum parser and no diff against the base documents. The delta record is modelled and validated and nothing produces one. An addendum in the package is read as one more document, not as a modification to the others. | Planned |
| Cross-document contradiction, DC-001 to DC-019 | No rule of that family exists in any pack and no module compares claims across documents. Quantity mismatch, product mismatch and conflicting responsibility assignment between two documents are not detected today. | Planned |
| Plan and drawing extraction | No module exists. There is no rasterization, no scale determination from a scale bar or title block, no tiled extraction and no review payload. Every spatial model the engine has ever run on was authored by hand. | Planned |
| Bill of materials as an input | The engine emits a bill of materials as supporting evidence. It does not read one in and check it against the documents. That direction is specified and unwritten. | Planned |
Refused by design, which is a different thing from unbuilt
Scanned documents
A PDF with no text layer is refused by name rather than silently ingested as empty. Route it through OCR and re-run. The refusal reports every unreadable file at once, so it costs one round trip and not one per file.
A package's own equipment schedule
Ingestion refuses any path under a schedule folder. A published equipment schedule is benchmark truth for that package, and a system that reads the answer key cannot be scored against it. There is no flag to override this. Moving the file is the override, and it leaves a trace.
The seam row is the one worth sitting with. The Division 08 to Division 28 boundary is the highest-value scope gap in this domain and it is what the engine is aimed at, and today it runs on one of its two evidence sources.
Rule governance: version, authority, and where to argue with it
A rule you cannot inspect is a rule you cannot dispute. Every family below is a JSON file in the engine repository, carrying its own version, its jurisdiction and the technical authority each rule cites, and every one is project and authority-having-jurisdiction overridable. A rule that reads a code section the way your jurisdiction does not is meant to be turned off and argued with.
| Family | Rule ids | Technical authority | Version | Last changed | Change history |
|---|---|---|---|---|---|
| Scope boundary scope_boundary |
SB-001, SB-002 | CSI MasterFormat Division 08 and Division 28 boundary; AIA A201 General Conditions 3.2 on the contractor's review of the contract documents; public contract code pre-bid clarification practice | 0.1.0 CA-DSA |
9 Aug 2026 853ee72 |
git log packs/scope_boundary.json |
| Division seam division_seam |
SB-020 to SB-024 | CSI MasterFormat Division 08 and Division 28 boundary; AIA A201 General Conditions 1.2.1 on complementary documents, 3.2 on pre-bid review, and 3.2.2 on reporting discovered discrepancies before proceeding | 0.1.0 CA-DSA |
9 Aug 2026 853ee72 |
git log packs/division_seam.json |
| Access control ac_k12_base |
AC-001, AC-010 to AC-014, AC-020 to AC-022, AC-030, plus validators V-001 to V-006 | NFPA 101 Life Safety Code 7.2.1.5; CBC 1010.2 and 1010.2.13; UL 294 Access Control System Units; ANSI/BHMA A156.31 Electric Strikes, Grade 1; NFPA 72 10.6 Secondary Power Supply; TIA-568 horizontal cabling distance; CSI 28 13 00, 08 41 13 and 08 71 00 | 0.1.0 CA-DSA |
26 Jul 2026 78030a9 |
git log packs/ac_k12_base.json |
| Video surveillance vs_k12_base |
VS-001, VS-002 | IEC 62676-4 Video surveillance systems, operational requirements; CSI 28 23 00 Video Surveillance; district security design standard | 0.1.0 CA-DSA |
26 Jul 2026 1b1f515 |
git log packs/vs_k12_base.json |
External review: not yet commissioned
No architect, code consultant, authority having jurisdiction or independent security engineer has reviewed these rule packs. The field is here, labeled and empty, rather than absent, because a governance table with no reviewer row reads as though the question was never asked. When a reviewer is engaged, their name, their standing and the date go in this position, and the rules they disputed go in the published failures.
Ids are permanent, and that is a constraint on us
A rule id is never renumbered and never reused. A superseded rule keeps its id, stays resolvable, and gains a pointer to the id that replaced it, because a published benchmark result cites rule ids and a renumbering would silently reattach an old number to new behaviour. That is why the empty blocks in the table above are reserved rather than left to be allocated later.
Every rule carries a rationale written for a practitioner and an authority citing a code section or a manufacturer document, and both print in the deliverable rather than living in the source. The rationale is written as if the reader is the engineer whose design it contradicts, because sometimes they are.
Read the published rulesWhat the numbers mean, and what they are not
Every item in the register carries an exposure band. Here is exactly what that band is.
It is an uncalibrated structural placeholder. Every band in every rule pack the project ships is marked as not calibrated, and that is not a hedge, it is a field in the data. None of them is derived from cost history. They exist so the ranking has a second key to sort on after severity, and that is the entire job they do.
Every rendering marks it. The uncalibrated marking is required on each place a band is printed, so a placeholder can never appear in the same typeface as a derived figure. A placeholder set in the same type as a real number is a confident output, and a confident output where the honest answer is a flag is the specific failure this project treats as unacceptable.
It is an order-of-magnitude planning figure. It is not a quotation, not an offer, not a price, and not derived from a takeoff. It is there to help you decide which three of thirty flags to chase this afternoon.
The deliverable carries no pricing at all. No unit costs, no labor rates, no extended totals. The bill of materials carries quantities and device identities and nothing with a currency on it, and that is enforced in the software rather than left to whoever writes the report.
No exposure figure appears anywhere on this website, for the same reason. Publishing a placeholder band next to a product claim would be using a number as evidence when the number is admitted to be a placeholder.
What it will not do
Not "has not got to yet." These are things the design refuses, and they will still be refused when the rest of it is finished.
It does not do takeoff.
It does not count devices off a drawing, because it does not read drawings. It does not produce your quantities. If you are looking for something that replaces the count, this is not it and will not become it.
It does not produce a quantity you should contract against.
The bill of materials it emits is supporting evidence for the register, not a purchase list. Every line names the rule that put it there so you can check it. Contracting against a number this tool produced, without checking it, is using it for something it was not built for.
It does not replace the estimator's review.
It does not decide what to bid, what to carry, or what to walk away from. It performs one specific review that gets skipped under deadline, which is the cross-check between documents, and it hands you a list of questions. A person answers them.
An empty result is not a clearance.
A short register means the rules that exist did not fire on the documents that were read. Given the table above, that covers a lot of ground the engine never went near. Read a short register as "these rules found nothing", never as "this package is clean."
The system is built to make that distinction where it can. A derived rule run against a package with no documents loaded raises an error rather than reporting zero findings, because zero findings and a clean package look identical in a report and are not the same thing. A validator run with no spatial model reports nothing rather than reporting a pass, because a validator that passed having examined nothing is indistinguishable from one that looked.
It does not arbitrate between two readings.
Where two independent readers disagree on a field, no value is emitted. Not a value with a warning next to it. There is deliberately no tiebreak, because a tiebreak reduces the flag count without reducing the ambiguity, and a shorter list reads as good news about the documents when it is only a change in how the software behaves.
The objection about verification
If you're going to have to go through and do the takeoff to verify what the AI spat out, what's even the point of using AI? Said out loud, repeatedly
The objection is correct about the thing it is aimed at. If a tool produces a count and you have to redo the count to trust it, you did the work twice and paid for software. Nobody should buy that.
It does not land on this tool for one reason, and the reason is a fact about what the output is, not an argument about model quality.
This is not takeoff. It is a review pass.
The output is not a quantity you check against your own quantity. It is a list of places in the documents where two sentences do not line up, each one naming the document, the page and the sentence. There is no number to verify because the engine is forbidden from producing one: the model is allowed to locate a sentence, transcribe it verbatim, classify it and match it, and it is never allowed to compute a count, pick a part number, produce a score, or rule on compliance. Every number in the output is computed by ordinary code from written rules.
What you do with a flag is not a takeoff. You open the cited page, read the sentence, and decide one of three things: it is a real gap and it becomes a pre-bid RFI, it is a real gap you will not get answered and it becomes a line on your qualifications and assumptions, or it is nothing and you close it. That last case is the false positive, and it is the cost.
What it costs, stated plainly
The cost is one open-the-page-and-read-it per flagged line, and it is paid by the estimator, in the week they have the least time. There is no version of this where the flag arrives pre-verified.
How much time that is depends entirely on how many lines come back, and that is the number nobody has. The only run anyone can point to is the repository's own hand-authored twelve-opening fixture, which produces a seven-line register with three items carrying an entity flagged for review. That is a synthetic fixture built to exercise the rules, not a real bid package, and quoting it as an expected workload would be dishonest. On a real package, the register length is unmeasured, which means the review cost is unmeasured too.
So the trade is not "software instead of review." It is: a fixed reading cost per flag, against the cross-check between documents that gets skipped first when the deadline compresses, because it is the slowest part and it usually finds nothing. If your process already runs that cross-check every time and finishes it, this tool has less to offer you than it does to someone whose process does not.
A pre-bid RFI is free. The same question asked after award is a negotiation.
The honest limit on this answer: it assumes the flags are worth reading, and that is exactly what has not been measured. If precision turns out to be poor, the reading cost per useful finding goes up and the objection starts landing. That number is the benchmark, and the benchmark does not exist.
The objection about absence
The AI can't tell you what's not there. Said out loud, repeatedly
True of a model reading one document. A model asked "is anything missing from this specification" has nothing to compare it against and will produce a plausible answer either way.
That is not how absence is found here. It is found by comparing two documents in code and naming both. The mechanism, rather than the conclusion:
How the engine finds an absence
- Two sets, built separately. For a given controlled opening, code assembles what the Division 08 sections say about its locking device, its power and its rough-in, and separately what the Division 28 sections say about the same three. Each entry carries the document, the section and the page it came from.
- The comparison is set logic, not a question to a model. Nothing asks a model whether the two agree. Code compares the sets, one opening and one aspect at a time.
- Five states, and four of them cost money. Both divisions carry the work, which is a duplicate buy. Both carry it and require different things, which is an argument on the day the door is hung. Each division assigns it to the other, which is the seam failure caught in the act. Neither division specifies it, which is the one where somebody eats it. And the fifth: the documents do not settle it.
- The finding names both documents. A "neither division specifies this" item cites the Division 08 section that was read and the Division 28 section that was read. Absence is reported as a relationship between two named, cited documents, which is a claim you can take into a pre-bid meeting and put in front of the person who wrote one of them.
- Undetermined is a separate finding from unassigned. They are not the same and collapsing them is what makes an estimator stop believing a register. "Neither section covers this" is a conclusion reached after reading both. "The documents do not settle it" is a flag that means look here. The engine has a distinct rule for the second one precisely so it never reports it as the first.
- No sections, no finding. If the package contains neither a Division 08 nor a Division 28 specification section, the engine emits nothing at all rather than reporting that neither division covers the work. A flag with no section heading to cite would carry no provenance, and an absence claim with nothing on both sides of it is not evidence.
The limit on this, stated with the mechanism rather than after it: this comparison reads two kinds of evidence, located scope statements and structured claims, and only the first is populated on any package the tool can currently produce. So the seam analysis runs today at half its designed evidence. That is the same row as in the table above, and it is the single largest gap between what this section describes and what runs.
Where to go from here
How it works
The design, in plain language. The AI reads and points, code decides, two independent readers must agree, and every line traces back to a page.
/security/Your documents
What happens to a bid package that gets processed, who receives text from it, where responses are cached, and what has to be disclosed before anything leaves the machine.
/The overview
One finding in full detail, the seam this is aimed at, the state of the build, and a longer list of what is not proven.
/diagrams/The same thing, drawn
Five diagrams with numbered points, including the pipeline with the unbuilt stages marked.
Blueprint Risk Engine. Every statement on this page about what is built, what is not built and what is refused traces to the project's architecture notes, its governing contract document, or a shipped rule pack. This page contains no benchmark result, no accuracy figure, no case study, no client, no testimonial and no pricing, because none of those exist to report. No bid specification text, district name, bid number or client material appears anywhere on this page.