About

Built from the problems that appear between the drawing, the BOM and the field.

I build Blueprint Risk Engine. It exists because the expensive failures in this trade are rarely hidden inside one document. They sit in the space between the door schedule, the hardware section and the security section, where each document is correct on its own. This page is who is building it, what the product refuses to do, and why.

Muhammad Sizar, PMP

Blueprint Risk Engine

Credential
Project Management Professional, Project Management Institute.
Focus
Division 08 ↔ Division 28 access-control scope, California public works.
Who is building this

The rules came from the work, not from a category study.

Blueprint Risk Engine is built by Muhammad Sizar, a PMP certified project professional. The experience behind it is direct rather than adjacent: project specifications, access-control hardware, bill-of-materials review, project management and project finance, IT, field implementation, and coordination between trades on California public works. That is a narrow set and it is exactly the set that produced the rules. Nearly every check in the engine started as something that had to be resolved after the fact, at an opening, with two documents open and neither of them wrong on its own.

  • Project specifications
  • Access-control hardware
  • Bill-of-materials review
  • Project management
  • Project finance
  • IT
  • Field implementation
  • Trade coordination, CA public works

The reason that list matters is narrow. Writing a specification, pricing it, managing it and then standing at the opening while it gets installed are four different views of the same document, and the gap this engine looks for is only visible from more than one of them. This was built by someone who has been in the room when the scope documentation and the field stop agreeing, and who had to answer for the difference.

The pattern that started it is the Division 08 to Division 28 boundary. A locking device, its power and its rough-in are described by the hardware section and by the security section. Sometimes both carry it, sometimes neither does, sometimes both do and they name different devices. Nothing looks wrong when either section is read alone. By the time anyone reconciles them the frame is prepped and the order is placed, and the question has changed from what was intended to who pays. A reviewer with a free week catches this. Bid week does not have a free week. That is a scheduling problem before it is anything else, and it is what the engine is pointed at.

The other half of the design comes from watching what an estimator does with a tool that overstates itself, which is stop opening it. So this one is built to be checked rather than believed. Every finding names the rule that fired and quotes the sentence that triggered it. Where the documents do not settle the question, the output is a flag with the uncertainty attached rather than a confident answer. The decisions behind that, including the ones that made the product less impressive, are written down in the repository, dated, and open to argument.

Technical questions get a technical answer. If you think a rule is wrong, say which one and why, and it will either be fixed or it will get a rationale worth reading.

The position

Six commitments, and what each one costs.

These are not values on a wall. Each one is a constraint written into the repository, each one is enforced by a test, and each one takes something away from the product in exchange for something else. The trade is stated so you can decide whether you agree with it.

01

A model may read. A model may not decide.

A model can extract text, classify a statement, locate a device on a sheet and match a schedule row to a specification paragraph. It never computes a quantity, assigns a part number, produces a risk score or determines that something is compliant. Those come from deterministic rules in version controlled files.

Costs
No ask-the-model fallback for a question the rules do not cover. Coverage grows one written rule at a time, which is slower than a chat box.
Buys
A finding that survives a design review. The answer to "says who" is a rule id and a quoted sentence, not a vendor.

Enforced by a test that walks the import graph and fails the build if a model client is reachable from the scoring code or the benchmark harness.

02

A finding without provenance is a bug, not a degraded finding.

Rule id, source document, page and the exact sentence, on every line that gets printed. Evidence is non-empty by constraint, so an item with nothing behind it cannot be constructed at all. It is not caught later in review; it never exists.

Costs
Things the engine might reasonably suspect, but cannot cite, are never reported at all.
Buys
You can open the PDF, land on the sentence, and disagree with the finding on the merits instead of on faith.
03

Where the documents do not settle it, produce a flag, not a guess.

An aluminum storefront whose stile depth cannot be verified gets a withheld locking device and a field-verification flag, not a plausible device. Where two independent readings of the same field disagree, no value is emitted at all and the disagreement itself is what reaches you. There is no arbitration and no averaging.

Costs
More flags, and flags take time to read. A guess would have looked cleaner in the register.
Buys
The register never quietly turns an unknown into a number that somebody then prices and lives with.
04

A rule that fires too broadly is worse than an absent rule.

Every rule ships with a positive test proving it fires on the case it was written for, and a near-miss test proving it stays silent on the case that only looks like it. A rule without the second test is not merged. The near miss is the harder test to write and it is the one that decides whether the output stays worth reading.

Costs
The rule count grows slowly. Checks that look obvious do not ship until the near miss is pinned down.
Buys
Output you still open in week three. A register that cries wolf costs the review time as well as the trust.
05

No number is published that has not been measured.

There is no accuracy figure on this site, no hours-saved figure, no dollars-saved figure and no pricing, because none of them has been measured. Exposure bands are order-of-magnitude planning figures, every band shipped today carries calibrated: false, and every rendering says so. They exist to give the ranking something to sort on, not to price anything.

Costs
No headline figure. The claims stop where the measurements stop.
Buys
Nothing here has to be walked back later, in front of the estimator who relied on it.

What we publish is measured. What isn't measured is labeled accordingly. The benchmark method was published before any result exists, so the method cannot be chosen after seeing the numbers.

What Blueprint does not detect
06

The same inputs produce the same output, byte for byte.

A run is pinned by its config hash and seed, and every model call is cached by content hash, with the vendor and model identifier in the key. The run replays identically later, including after a vendor retires the model that produced it.

Costs
No silent upgrade to whatever model is newest this month, and no orchestration that decides its own control flow.
Buys
A run that can be replayed. A run that cannot be replayed cannot be benchmarked, and an engine that cannot be benchmarked is asking for trust it has not earned.
Why this is built in the open

A finding is only worth something if somebody can check it.

The rules are published in full, with their conditions, rationale and cited authority. So are the limitations, the current build state and the benchmark method, including the parts that are not finished and the checks that are allocated but not written.

The reasoning is practical rather than principled. A preconstruction tool that cannot be audited cannot be defended in a scope dispute. When a finding turns into an RFI, a qualification or an exclusion, somebody on the other side asks where it came from, and the answer has to be a document, a page and a named rule. If the answer is that a system said so, the finding is worth nothing at the moment it matters most. Publishing the checks is not generosity. It is the only form in which the output is usable.

Build state is at status. The measurement method is at benchmark, where there is a method and no result.

Where it is going

What is next, without dates attached to it.

No roadmap is published with dates, because a date published now would be a claim about work nobody has scheduled. What follows is the shape of the gap and what would change the order of the work.

Jurisdiction

The seam and scope logic is jurisdiction independent: two sections assigning the same opening to each other is the same finding anywhere. The access control and video packs are not. They are written against California DSA reviewed K-12 public works, and all four shipped packs declare that jurisdiction. Commercial packs are not written yet, and running these ones on a commercial package would cite authority that does not govern it.

The benchmark

The dataset is being assembled: packages where the change orders actually issued are a matter of public record, so a pre-bid flag can be scored against a realized outcome instead of against an opinion. Until that exists there is no precision or recall figure, and there will not be one. The scoring harness is written and the method is published; the adjudicated cases are the missing part.

Rule coverage

Rule id blocks are allocated in advance and never renumbered, because published results cite them. Several blocks are reserved and empty today, and the build status page names which ones. Coverage, calibrated exposure bands and a published benchmark are the binding constraints, in that order. Model capability is not one of them.

What would change the priority is what pilot participants report. Which findings were useful, which were noise, and what the engine missed entirely. If the seam findings hold up and the noise turns out to be somewhere else, the next work goes where the noise is, not where a plan written in advance said it should go. That is the whole reason the pilot exists, and it is why there is no pricing yet.

Run it against a real package and tell me where it is wrong.

A limited number of electronic security contractors and estimators are being taken on as design partners. The useful part of a pilot is the feedback, not the demo.