← Omkar Khadamkar
Enterprise security platform · 2022 – 2026

Running a design system an AI writes against

I led the design system for an enterprise security platform. Its highest-volume consumer stopped being a person some time ago — and almost everything about how you govern a system changes once that is true. Including how you answer the question an engineering leader will actually ask: what did this buy that a component library off the shelf would not have?

On-system accuracy ~25% → ~90% Correction per screen ~2 hrs → 15 min Design time per screen −22% Designers on the system 5, none reporting to me
In plain terms

A design system is a set of rules for how a product should look. People used to be the only ones reading them. Now an AI reads them too, and it builds far more screens than any of us do.

An AI reads differently from a person. It never asks what you meant, it never notices that a rule is vague, and it never leaves a blank. So the documentation had to change, the order we build things in had to change, and we had to start measuring whether the system actually held.

Getting that right turned out to be worth roughly a fifth of the time it takes to produce a screen.

Contents · 8 sections · ~9 min
  1. 01What changed
  2. 02What it bought
  3. 03Governance you cannot ignore
  4. 04Foundations before consumers
  5. 05Foundations that carry meaning
  6. 06The generation layer
  7. 07The unglamorous half
  8. 08What is not here
01What changed

A design system has always had two audiences: the designer opening the Figma library, and the engineer reading the spec. A third one arrived, and it consumes documentation faster than both combined.

The difference is not that the model is worse. It is that a person and a model fail differently. Hand a designer an under-specified rule and they ask, or they copy a neighbouring component, or they leave it and flag it in review. Hand the same rule to a model and it produces a confident, plausible, wrong answer — and then repeats that answer consistently enough to look like a decision.

Which means the two things a system is normally judged on — is it beautiful, is it adopted — are joined by a third: does it survive being generated?

Everything below follows from taking that question seriously.

02What it bought

Back to the engineering leader’s question, because it is the one that decides whether any of the rest gets funded. My honest answer to it, for far too long, was directional — faster, more consistent, fewer rounds. Directional is not an answer. Someone deciding where two headcount go needs a number with a method behind it, and “let me get back to you on that” loses the room before you have finished saying it.

So: the number, and the counterfactual it is measured against.

~25%
On-system, with the
component library alone
~90%
On-system, once the rules
were machine-readable
~2 hrs → 15 min
Correction per screen,
before and after
−22%
Total design time
per screen

The first figure is the whole argument. A component library on its own — ours, Material, Tailwind, any of them — got generated UI about a quarter of the way onto the system. The library is the part you can buy. The remaining sixty-five points came from writing the rules so a machine could not misread them — exact values rather than relative descriptions, component status encoded where the work happens, a style layer bound to the production tokens rather than restating them. That part is not purchasable, because it has to describe your product rather than somebody else’s.

ON-SYSTEM ACCURACY OF GENERATED UI 0 25% 90% 100% 25 +65 COMPONENT LIBRARY THE PART YOU CAN BUY MACHINE-READABLE RULES THE PART YOU HAVE TO WRITE STILL OFF
A component library on its own — ours, Material, Tailwind, any of them — got generated UI about a quarter of the way onto the system. The remaining sixty-five points are not purchasable, because they have to describe your product rather than somebody else’s.

The step that moved is correction: pulling generated output back onto the system by hand. It ran around two hours a screen and now runs about fifteen minutes. Correction sits inside a screen that takes roughly eight hours end to end, so the saving is a little over a fifth of total design time.

The accuracy percentage and the hours saved are not two findings. They are one thing measured twice — at 25% on-system a designer rebuilds three-quarters of it by hand, and at 90% they check it.

Published gains for mature design systems run 34–50% on design productivity. Mine sits below that band and I would rather it did. An earlier version of this arithmetic used a three-hour correction average and landed on 34% — precisely the benchmark floor, which is too convenient a place to land. Revising my own average down to two hours cut the headline figure by a third and changed none of the conclusions that rested on it. A number that survives being wrong is worth more than a larger one that doesn’t.

On the engineering side there is one measured instance and no rollout. A page went from Figma to a running implementation in about an hour, once its components were mapped to code, against a conventional estimate of two to three developer-days. That is a single page; the mappings were mine; and the hour covers design to correct composition, not design to done. I would say all three of those things before quoting the first one.

The figure worth carrying into a different organisation is not my total — that depends on a former employer’s volume, which is theirs and not mine to publish. It is the unit underneath it: about 1.75 hours of design time returned on every screen, before anything engineering-side. Multiply that by your own volume rather than trusting mine.

ONE SCREEN, END TO END BEFORE 8 hrs AFTER 6.25 hrs EVERYTHING ELSE · UNCHANGED 2 HRS CORRECTION 15 MIN 1.75 HRS RETURNED PER SCREEN — A LITTLE OVER A FIFTH OF TOTAL DESIGN TIME
Only one step moved. Correction — pulling generated output back onto the system by hand — is the whole of the saving; the rest of the screen costs what it always did.

And what the arithmetic will not support: any claim about organisational adoption, engineering hours saved at scale, or what happens once the people prompting the system are not the people who built it. Those want measuring, not modelling.

The figures have a sensitivity table behind them, and payback stays under six months in every row — including the one where all three assumptions are cut at the same time. An earlier version of this model, a third larger, paid back in the same window. That the answer barely moves when the inputs do is worth more than the answer.

03Governance you cannot ignore

Most design systems keep governance in a document nobody opens. The status of a component — approved, buggy, deprecated, do-not-touch — lives in a wiki, and the Figma file quietly disagrees with it.

I moved the lifecycle into the file structure itself, so a component’s status is visible before you can use it:

✅ApprovedBuilt, reviewed, specified. Safe to consume.Use freely
🐞Known bugsReal, in use, and carrying defects that are already logged.Use with care
🚫Main components — do not touchThe sources everything else instances from.Locked
🧪ExperimentalBeing explored. May change shape or disappear.Do not ship
❌RetiredSuperseded. Kept only so old work stays readable.Do not consume

It looks almost too simple to count as governance. That is the point. A rule that lives where the work happens gets followed; a rule that lives in a document competes with the work for attention and loses.

It matters more once a model is involved, because a model has no way to ask whether a component is safe to use. It will happily instance a deprecated component or an experiment if nothing in its context says otherwise. Encoding status structurally means the answer travels with the thing.

Governance that lives in a separate document is governance you are choosing not to enforce.

04Foundations before consumers

The most expensive mistake available in a design system is authoring a shared thing second.

A concrete case. A container mechanism — a slot that holds consumer-supplied content — was going to be shared by three components. The tempting move was to build the first component that needed it and let the slot emerge from that work.

AUTHORED IN THE WRONG ORDER Consumer 1 Shared contract Consumer 2 bends Consumer 3 bends shaped around consumer 1 alone AUTHORED IN THE RIGHT ORDER Shared contract Consumer 1 Consumer 2 Consumer 3 each one a short spec
Author the consumer first and you are authoring against a moving foundation. The rewrite is never in one place.

So the slot was pulled out and specified on its own first: what a slot is, what a consumer may inject, and where the boundary sits between container-owned chrome and consumer-owned content. Each component then became a short spec that says “a slot container, positioned like this, dismissed like that.”

The same discipline catches a subtler failure I have come to think of as a borrowed token doing double duty. A token gets reused for a second purpose because the value happens to match. The moment either purpose needs to change, you cannot move it without breaking the other, and the only fix is to unpick both.

Values matching is not the same as values meaning the same thing.

05Foundations that carry meaning

A type ramp is where a system either states its domain or hides it. Ours carries the usual semantic scale — headings, body, meta — and then a second scale that exists purely because of what the product is:

Large stat1,284
Medium stat98.6%
Small stat42
Mini stat7 open
Name / valueRegion  eu-west-1

A security platform is mostly numbers under pressure — counts, percentages, thresholds, deltas. Treating stats as body text and scaling them by eye is how dashboards drift into inconsistency, so they get their own ramp, paired with a name/value convention for the label-and-figure pattern that appears in every panel.

The link scale earns its place the same way. Links exist inline in a sentence, inline in a dense table, and standing alone as a button-like action. Those are three different jobs, so they are three defined styles rather than one style plus judgement — because judgement is exactly what a generator does not have.

06The generation layer

On top of the system sits the part that makes it machine-consumable: component guidelines written so a model cannot misread them, specs carrying exact values rather than relative descriptions, and a style layer bound to the production token file rather than restating it.

The test is blunt. Regenerate a page from the guidelines alone, name no fix in the prompt, and count what lands.

25 → 65 → 90%
Library alone, first draft,
after testing and rewriting
4–5 → 1–2
Regeneration rounds
to ship-quality
5
Designers on
the system
All
Product areas
covered

Those figures are my own measurements on our system, not instrumented telemetry, and I would describe them as approximate. The method behind them is public and reproducible — I rebuilt it from scratch on a generic system so it could be checked without anything proprietary in it, and in doing so found that one of my own published claims about why it works was wrong.

Field test 01 →   Field test 02 →

07The unglamorous half

None of the above matters if the system is not used by people who did not build it. That is the half with no craft in it, and it is the half that decides whether the rest survives.

111
Component-to-code
mappings, in a month
17
Requirement stories on
one spec, three teams
6
In-product guides
shipped instrumented
427
Unique users on the
lead guide, six weeks

Five designers worked on and against the system, across every product area on the platform. So the day-to-day was less about designing components than about the things that decide whether a system holds: reviewing contributions, keeping the library and production honest with each other, deciding what was worth standardising and what should stay local, and coaching the people whose work depended on it.

Three of those had outcomes I can point at. One designer went from occasional contributor to owning the icon set outright. One was promoted into wider scope. And a guidance practice I had built was picked up by leads senior to me — which is the version of teaching that is hardest to arrange, because nobody above you has to take it. How that works without a reporting line →

Adoption is built, not requested. Engineering could consume the system directly because 111 components were mapped to their code equivalents over about a month, so pasting a design link into an editor returned the real component — states, edge cases and accessibility already in it — rather than a picture to reimplement. Measurement worked the same way: one instrumentation spec pattern carried across seventeen requirement stories in three teams, so what shipped could be read the same way regardless of who built it.

Then you find out whether anyone used it. Six in-product guides shipped instrumented, and the lead one reached 427 unique users at 35% engagement in six weeks. Two things in that reporting are worth more than the number. One guide came back with zero engaged users and stayed in the report. And one metric read close to 100% click-through, which I footnoted so the room would stop believing it. The figure that flatters you is the one to check first.

The measurable version of adoption is not a dashboard. It is whether a new screen can be built without anyone inventing a value — and whether, when a model builds it instead, the result still passes an audit.

What I cannot tell you is whether it held. I left in August 2026, and the real test of a system is whether it outlives the person who wrote it down. That one is running without me.

08What is not here

The components, screens, token names and product surfaces belong to a former employer and are not shown. What is described above is architecture and method: how the lifecycle is encoded, why shared contracts are authored before their consumers, how a token scale carries domain meaning, and what changes once a model is a primary reader of your documentation.

The measurement method is public in full, rebuilt on a generic system so it can be inspected and re-run by anyone: github.com/omkarpkh/design-system-drift-lab

One last thing, and it works against everything above. A widely cited industry estimate puts roughly 45% of shipped features in the never-used column. So this makes a pipeline cheaper whose largest loss sits upstream of it. Making the wrong screen 22% cheaper is still the wrong screen — worth saying out loud before anyone treats a design system as a productivity programme.

I am happy to go deeper in conversation on any of it.