Every framework arrives with an assessment, and every assessment finds you deficient. This is not an accident of method. A maturity model published by someone with something to sell scores you against the top of the ladder, and the distance between you and the top is the sales pipeline. After nearly three decades around these instruments I can count on one hand the assessments that ended with "you are fine, spend your money elsewhere".
The previous article ended with a claim: an honest instrument has to be able to say "this is fine" as often as it says "this is the gap", or it is not an instrument, it is marketing. This article is about the instrument itself. The diagnostic from Beyond the BOM is now an interactive page on the book's site. You can click through it in ten minutes; done properly, with the right people in the room, it takes thirty. Everything you enter stays in your browser, and it is built to be argued with.
What it actually asks
Eight sections, one for each primitive and one for the federation pattern that carries them. Each section asks the same underlying question in a different domain: does this exist as managed engineering data, or is it inferred, remembered, and reconstructed? System: is a system a first-class object with an owner and a lifecycle, or an implicit hierarchy read out of the BOM tree? Interface: a managed contract, or a link inferred from part adjacency? Verification: a continuous capability with lineage, or a series of test reports attached to release events? Configuration history: can you answer "what was true when" as a query, or is the answer a reconstruction project?
Each section offers five levels, from absent to continuous. The levels are not grades to aspire to. They are descriptions, and your job is to find the one that reads like an ordinary Tuesday in your organisation, not like your last management presentation. Notice also what the sections do not ask: nothing about which vendor you run, how modern your tools are, or how large you are. The instrument measures the data model, because that is where the problem lives.
The dashed line is the point
Scoring eight sections produces a radar chart. On its own that is decoration. The reason the diagnostic works differently from the assessments I complained about above is the second polygon: before interpreting anything, the page asks which of the four scenarios from the previous article is closest to your product, and it draws your scenario's demand line as a dashed shape on the same radar. The reading is the gap between your solid shape and that dashed line. The outer ring belongs to the regulated system alone, and if you do not build one, it is not your target and never was.
The instrument is also suspicious in the other direction. Score yourself at the top level across all eight sections and it does not congratulate you. It flags the result as optimistic, because in my experience the organisation that operates every primitive continuously does not exist, and a perfect score usually measures the scorer, not the data model.
A worked reading
Take the machinery company from the previous article, the one that started selling service contracts on a serialised fleet. Suppose it scores itself honestly: system 3, interface 3, space 2, verification 3, configuration history 2, decision 2, judgment 1, federation 2.
Read against the mechatronic demand line, this radar is comfortable. At or above the dashed line nearly everywhere. That is not an illusion; it is a correct statement about the company this organisation used to be.
Read against the industrial demand line, the same radar changes meaning. Configuration history sits at 2 against a demand of 4, the widest gap on the chart. Judgment sits at 1 against 3. Space, decision, and federation each run one level short. Nothing on the chart moved. The business moved, from products to a serviced fleet, and the dashed line moved with it.
Look at which gaps are widest: configuration history and judgment. The previous article argued the primitives arrive in a predictable order, and that configuration history arrives with serialisation while judgment arrives when product life outgrows the memory of the people. Those are exactly the two boundaries this company just crossed. The radar is not telling you where you are weak in general. It is showing you the front line, the place where your drivers most recently crossed and your data model has not yet noticed.
Why as-maintained sits at level four
The demand lines mostly ask for level 3, managed objects in the working domain, at the industrial tier. One cell is deliberately different: configuration history demands level 4 already there, and since this is the calibration most worth arguing with, it deserves its defence in public.
Level 3 for configuration history reads respectably: per-system records well maintained, cross-system reconstruction documented and repeatable. For a company that ships products, that is enough. For a company that services a serialised fleet, it is not, because the service business asks per-instance questions at commercial frequency. Which delivered units contain the substituted component. What was this unit's configuration when the failure occurred. Which units does the retrofit quote cover. A repeatable reconstruction answers each of these in days of engineering effort, and the warranty case, the recall scope, and the quote deadline do not wait days. At this tier the as-maintained configuration must be queryable, not reconstructable, and that is the definition of level 4.
That is the argument. If your serviced fleet runs fine on repeatable reconstruction, I would like to hear how, and the comments are open for it.
Above the line is also a finding
An instrument that can only point upward is a ratchet, not a gauge. Scoring above your dashed line is a finding too: it means you are carrying formalisation your drivers do not demand, and formalisation has a cost, a permanent one, in discipline and maintenance. Judgment objects with validity triggers for a three-year product built by one team is not rigour. It is governance nobody will attend, and most governance nobody attends started life as somebody's above-target ambition. The right depth is the lowest one the drivers allow, and an honest instrument says so to your face.
One consequence for the team that runs it
Do not run it alone. A radar scored by one person measures that person's optimism. Run it with a designer, someone from service, and someone from quality in the same room, and treat the disagreements as the primary output. When the design engineer says configuration history is level 3 and the service manager laughs, that laugh contains more diagnostic information than the chart does. Each section has a notes field for capturing exactly this, and like everything else on the page it stays in your browser. Bring the radar and the disagreements to your next architecture discussion. It is probably a better agenda than the one the vendor brought.
One more thing
The diagnostic condenses Beyond the BOM: A Federation Architecture for Complex Engineering Data into eight questions. The book develops each answer at the depth the regulated system demands: the primitives, the federation layer, the roles, and the operating cadence, across 395 pages now on Amazon. The earlier articles in this series argued that the digital thread, the AI copilot, the knowledge-retention programme, and the replacement fantasy all fail on the same missing substrate, and the previous one mapped which pieces of it your product actually needs. This one hands you the tape measure. What you do with the reading is, as it should be, your decision and not mine.