Don't trust a black box.
Inspect the rules.
Most software in this category asks you to believe a number. Blueprint publishes the checks, the gaps, the build state, the run artifacts and the measurement method, including the parts that are not finished.
What we publish is measured. What isn't measured is labeled accordingly.
Six pages, all of them one click from here.
Published rules
Every check the engine performs, with conditions, rationale and cited authority, taken verbatim from the rule packs. Nineteen rules across four packs and six whole-system validators, each one reproduced in the words the pack file uses rather than paraphrased into marketing copy. The id blocks that are allocated and still empty are listed too, so you can see what is reserved and not yet written.
Read the rules 02Limitations
What Blueprint does not detect. What it has never read. What its numbers are and what they are not. It includes a real miss, recorded rather than tuned away: on the first project manual the scope detector was run against, it walked past a page that assigned owner-furnished scope in ordinary English with none of the marker phrases in it.
Read the limitations 03Build status
What is implemented, what is partial, and what is not built at all, phase by phase, with run results rather than estimates. It also lists the defects the last session found in code that was already reporting itself as working, because a green test suite is evidence that the tests ran, not that the code is right.
Read the build status 04Benchmark
The measurement method, published before any result exists. What will be counted, how a human adjudicates a flag against a change order actually issued, and the rules that stop the benchmark from flattering the engine. There is no precision figure and no recall figure on it, and that is the point of reading it now.
Read the method 05Technical architecture
The page for the engineer your estimator forwards the link to. The real command and the three files a run writes, the config hash and the byte-identical replay it proves, the schema a finding has to satisfy, and the import-graph test that stops a model client from ever being reachable from the scoring code or the benchmark harness. It is the part of the record that is checkable by reading rather than by believing.
Read the architecture 06Pipeline concept demo
The intended production workflow, end to end, labelled a concept because that is what it is: the walk-through is a design, not a deployed application. What sits inside it is not a mock-up. The findings shown are the real fixture-run output, with the rule ids, the citations and the counts the engine actually produced. Read it to see the shape of the workflow, and read the build states on it to see which parts of that shape run today.
Walk the conceptEach of those pages gives you something to argue with.
None of them is a trust exercise. Each one hands you a specific thing you can check, reject, or plan around without asking anyone here a question.
A rule you can read is a rule you can reject
Some of the nineteen will be checks you already run. Some will be checks you do not. A few you will think are wrong for your jurisdiction, and you will be right, because every rule is project-overridable and the packs shipped today are written against California DSA reviewed K-12 work. All three reactions are useful and none of them requires taking a vendor's word for anything.
A limitation you can read is a limitation you can plan around
A tool that tells you it never read the addendum is a tool you can hand the addendum to. A tool that stays quiet about it is one you find out about after award. The list is kept current because a stale limitations page is worse than no limitations page.
A method published before its results could not be chosen to flatter them
A benchmark whose definitions were settled after the first run is not a benchmark, it is a selection: thresholds move, awkward cases get excluded, and the metric that looked worst turns out to have been the wrong metric all along. The only defence is to write the method down first and be held to it.
Read the pages, then look at what the engine actually printed.
The demo is a sample risk register from a fixture run: the findings, the rule that fired on each one, and the sentence in the documents that triggered it. It is the shortest way to decide whether any of this is worth a package of yours.