Two numbers, plotted, never merged: how likely a product is to be non-compliant, and how dangerous it would be if it is
Every regulated product-type instance an agency oversees ends up as one point on a grid like the one below. Nothing downstream multiplies the two axes together into a single "risk score" — an agency works its queue by scanning the plane, not by reading down a sorted list. The highlighted case is built up step by step across the rest of this page.
The color wash behind each cell is a reading aid, not a stored number — nothing here averages the two axes into a value that gets saved or sorted. These five points follow the same structure and scoring as a real assessment run, with illustrative specifics.
Your rules, fixed first
Your taxonomy, your objectives, your risk-criteria matrix — approved once, up front. No model touches or redefines them; they're simply what gets scored against.
Drafts every score, checked automatically
Two model passes assess likelihood and severity for every product, each producing a factor score or a scenario with the one-line evidence that justifies it. Structural checks, litmus tests, and deterministic recomputation stand between a draft and a stored result — before you ever see it.
You decide, always
The assessment is drafted for you, with evidence behind every score. Nothing here acts on your behalf — you're the one who reads the plot, weighs the case, and decides what happens next.
Every regulator's world is fixed before a model ever runs
The taxonomy — categories, families, individual product types — is reference data a regulator approves once, up front. No model call touches it. It's simply what the AI layer is scoring against.
| Regulated category | Families | Product types |
|---|---|---|
| Construction materials & building products | 5 | 12 |
| Chemicals & chemical products | 5 | 15 |
| Electrical & electronic products | 3 | 7 |
| Cosmetics & personal care products | 1 | 6 |
Cement sits in Mineral construction materials — one of 5 families under Construction materials & building products, one of 12 product types in that category, and one of 40 shown here. A regulator's full taxonomy might run to, for example, 140 product types across a dozen-odd categories; food, beverages, and agricultural inputs would make up the rest and sit outside this page's scope. This case is built up step by step from Section 07 onward.
What "consequence" means is agreed before the first case
Each dimension of harm comes with four pre-written levels, Low through Very High, approved by the regulator ahead of time. A model drafts scenarios later — it never gets to define what "High" means. The two ladders below are reproduced verbatim from an approved risk-criteria matrix.
Consumer Health & Safety
Fair Competition
A fixed list of "why," scored per case
Behind the two axes sit two catalogues: probability factors (why non-compliance might happen) and severity factors (why it would matter). Both are regulator-owned and fixed. What's not fixed is which factors apply to a given product — that's the model's first job, next section. Codes and wording below follow the structure and style of a regulator's own factor lists.
Probability factors — illustrative selection
Severity factors — 6 dimensions, 39 factors (8 shown)
Consequence gets drafted by two passes, not one
Every assessed instance goes through two model calls before a severity number exists. One covers ground; the other goes deep on what that ground implies.
Scores every factor
Checks every severity factor in the library against this product type in general, and may propose factors genuinely missing from the list — a real gap, not a hallucination to be trusted blindly (Section 06 covers how that gets checked).
Ungrounded — reasons from the fixed factor libraryDrafts the scenario
Takes Agent 1's applicable factors as given input and constructs plausible non-compliance scenarios, scored against the regulator's own Low–Very High matrix.
Ungrounded — same reasoning basis as Agent 1A third, separate call scores probability — and unlike these two, it's grounded in live web search with real citations. See Section 07.
Four guardrails between a draft and a stored number
"AI-supported" doesn't mean the model's output is taken on faith. Every draft passes through checks that are literal code, not further prompting.
The litmus test
A factor only counts if the harm is contingent on the compliance gap itself — not on the product's inherent nature.
Severity anchors
The count of triggered factors in a dimension sets a floor on that dimension's level — thresholds derived from the regulator's own factor-list size, never a fixed table.
Structural validation
Every response is checked against a strict schema — required fields, valid factor codes, and a value/score consistency rule — before it's accepted.
Deterministic recomputation
Aggregation, anchor-floor checks, and every continuous total are recalculated in Python from the model's raw output.
Cement: how likely is non-compliance?
Every probability factor is scored 0–1 for this specific product type, each with a one-line evidence note. Unlike Agent 1 and Agent 2, this call is grounded in live web search. This is a real run: the six factors below are a sample of the 22 actually scored.
0.49 isn't the average of the six factors above. It's what Guardrail 04 recomputed in Python from all 22 factors in the library — including several, like MS6 and TS6 here, that correctly score near zero because they simply don't apply to cement. Showing only the high scorers would have overstated it. The index never touches consequence — it judges how urgently to look here at all, a separate question from what happens if it goes wrong.
Cement: how bad, if non-compliance reaches the market?
Consequence is established twice, independently, and cross-checked — Agent 1's factor pass, then Agent 2's scenario, each held to the guardrails from Section 06. This is a real assessment: the factors and scenario below are reproduced from an actual run.
Consequence isn't one number even within itself
The regulator's ordinal matrix level reflects the single worst scenario Agent 2 could construct — for Cement, three of the six dimensions reach Very High this way. The continuous consequence index reflects how broadly the severity factors apply overall, across all 39 factors in the library. They diverge, and both get reported rather than reconciled into one.
Neither number overrides the other. Cement's continuous index (0.72) sits in the "High" band on its own, while its ordinal level is a full step higher at "Very High" — driven by what the worst plausible scenario looks like, not by how many factors apply on average. A regulator needs to see both to know which is true here.
Back to the grid — risk-based, never blended
Probability 0.49. Consequence 0.72. Neither number changes the other, and neither gets multiplied into a single "risk score." Plotted together against four other real assessments from the same regulator's queue, Cement sits high on both axes — exactly the information a blended score would have thrown away.
The color wash behind each cell is a reading aid, not a stored number — nothing here averages the two axes into a value that gets saved or sorted.
Wooden Clothes Pegs sits low on both axes and can wait for a routine cycle. Cement and Insulated Electrical Cable land close together in the upper-right — not because they share a hazard, but because two very different patterns each push both axes up independently: weak testing-and-standards coverage for cement, a fragmented and unregistered manufacturing base for cable.
Not the same as a "most severe scenario" ranking. A separate ranking, computed only within one regulator's results, orders products by how bad their single worst drafted scenario is — useful for picking representative cases when choosing what to test, not for deciding what to check first. Targeting runs off the continuous, factor-derived numbers plotted here instead.
Every recommended test traces back to a scenario above
The scenario, not a generic checklist, decides what the inspection floor should actually check for. Both tests below are real recommendations from the same Cement assessment.
Three actors, not one AI black box
Fixed before any model runs
Every regulator's taxonomy, objectives, and factor library is approved reference data. Nothing here carries over between agencies, and nothing is invented per case.
Drafts, then gets checked
A breadth pass, a depth pass, and a grounded probability call — each producing a score or a scenario with the sentence that justifies it, some backed by live search citations. Structural validation, litmus tests, severity anchors, and deterministic recomputation stand between a draft and a stored number.
Decides, always
Probability and consequence stay separate, always, and neither gets acted on automatically. The assessment is drafted for you — you're the one who reads it and decides what happens next.