Skip to content
Insights

AI makes good design systems faster and bad ones messier
Is your design system ready for AI? Preparing your product for a new way of building

Most organisations treat their design system as a design and engineering concern: a component library, a Figma file, a Storybook. It rarely comes up when the business talks about strategy.

That is starting to change. Products now span more platforms, brands, and teams than ever, and AI is impacting how quickly software gets built. Both trends make it even more valuable to have a clear, shared definition of how your product should look and behave.

A few years ago, I read Alla Kholmatova's Design Systems, and one idea has stayed with me: a design system is more than a collection of reusable components. It's a shared language.[1] We've already seen that language mature, from describing what things look like (#0055FF) to describing what design decisions mean (color.action.primary). What has changed is how many things now need to speak it, including machines.

When that language is unclear, it costs the organisation in duplicated work, inconsistent experiences across platforms, and stalled migrations. AI amplifies whichever situation - and whatever mess - you're in. Coding agents and design tools can already read structured design context.[2] So give them a well-defined system and they’ll accelerate good work. Give them a fragmented one and they’ll produce inconsistency faster.

As the cost of producing code falls, the implementation itself is now secondary to the definition behind it. That includes the intent and evidence that tells us whether an implementation is correct. Organisations that make their product language explicit and governed, whilst also being machine-readable, can adapt to new environments and tools without starting again each time. Those that don't risk paying for the same decisions over and over.

Much of the industry conversation is now about orchestration: directing agents rather than writing code. But orchestration is often treated as the system itself, when it is really only one layer. Without the right harness, context, constraints, and definition of quality, coordinating more agents simply creates more output to review. Orchestration is only as good as what it’s directed against. Agents need a clear definition of what good looks like, and for product teams, much of that definition lives in the design system.

So how do you build a design system that both humans and AI can use?

I’m sharing some of what we've learned from working with organisations in complex and regulated environments to get their design systems agent-ready. It's a long read, so here's the route through it. We start with why the design system no longer lives in one place and how that creates drift. Next we look at how tokens and component contracts make intent explicit for both people and machines, and why governance and feedback loops keep that intent aligned. Then comes the practical reality of engineering standards, existing codebases and migration. Finally, we cover what cheaper implementation and AI mean for where the real value of a design system lies.

Contents

The design system doesn’t live in one place

Across different products and organisations, we’ve seen the modern design system become increasingly distributed. Ask a designer where the design system lives and the answer might be Figma. Ask an engineer and it might be the component library, or Storybook. Someone responsible for governance might send you to Zeroheight or another documentation platform. They’re probably all partly right – and the applications themselves tell us how the system is actually being used.

Each tool contains a view or implementation of the underlying product language. No single one contains the complete design system.

Fragmentation creates drift

The specialised tools aren’t really the problem, it’s what happens in the spaces between them.

A designer updates a primary-action token in Figma but production does not change. Or engineering adds a loading state that never makes it back into design. Small gaps like these compound across a product organisation - and drift is not only visual. Values, component states, behaviour, documentation, and local product variants can all diverge from the shared system.

This is why I’m increasingly cautious about the phrase ‘single source of truth’. It’s tempting to solve the problem by declaring that Figma is the source of truth, or that everything should ultimately live in a Git repository, but the boundaries aren’t that clean.

Production code is authoritative about what gets executed. Figma may be authoritative for aspects of the design expression. A component contract might define the intended behaviour - and testing monitors the continuity of that behaviour.

What is authoritative for each kind of decision, and how do we keep all of its consumers aligned?

Instead of trying to force every part of a design system into one tool, we can start thinking about how information moves between those tools – and how we know when they disagree.

From shared values to shared meaning

Design tokens aren’t new, but what we now have is an ability to describe them in a way that isn’t tied to a particular design tool or engineering platform.

The Design Tokens Community Group reached an important milestone in October 2025 with the first stable version of its specification, 2025.10, which the group describes as a production-ready, vendor-neutral format for sharing design decisions across tools and platforms.[3]

The value comes from interoperability rather than the JSON format itself.

Instead of independently defining the concept of a primary action colour in Figma, Swift, Kotlin, and CSS, we can increasingly start from the same semantic decision and transform it into whatever representation makes sense for each consumer. They need to share intent rather than the implementation.

That makes naming more important, but it doesn’t mean forcing identical syntax everywhere. ‘color.action.primary’ might become ‘color/action/primary’ in Figma, ‘Color.actionPrimary’ in Swift and ‘--color-action-primary’ on the web. Consistency of meaning matters more than consistency of syntax. Clear mappings allow each representation to refer to the same governed concept.

Tokens only take us so far. A token can identify the colour used for a destructive action, but it cannot explain when that action is appropriate. It also cannot define the states a component supports or the behavioural and accessibility constraints that apply. That requires a fuller description of the decision.

From components to contracts

A component already has an implicit contract. If we take the example of a button, its variants, states, inputs, behaviour, accessibility, and usage rules form that contract.

Today that knowledge is often scattered across Figma, code, Storybook, documentation, and the people who know the system well. Humans can navigate that ambiguity by checking several sources and asking questions, but machines need that knowledge to be explicit.

This is where we've been exploring the idea of a component contract: taking that implicit knowledge and writing it down in a shared, structured form. A contract goes beyond what a component can do. It captures why it exists, when it should be used and the rules it must always follow.

If tokens describe the vocabulary of a design language, contracts begin to describe its grammar.

Knowing that Button exists is useful; the next step is understanding when and how its variants should be used. This gets us much closer to the actual product decision.

A language for people and machines

Structured definitions should complement human documentation, not replace it. People still need explanation and context, while machines need structured definitions they can discover and reason about.

Your team might encounter a component through Figma, Storybook, or documentation; an agent through metadata, an API, MCP, or the repository. We think of them as different views of the same governed knowledge versus two independently maintained versions of the system.

If an agent implements a finished design, many product decisions are already made. If it composes an interface for a task, it is participating in those decisions. It therefore needs to understand not just what components exist, but their purpose, composition patterns, and constraints.

Making a design system machine-readable helps an agent generate a better button and gives it enough understanding of the product language to decide which button to use – or whether a button is appropriate at all.

A design system has always been a language for humans. Making that language machine-readable gives AI the opportunity to participate without inventing a language of its own.

Governance is part of the architecture

Governance can sound administrative, but it becomes an architectural concern once a system serves several products and implementations. A token rename, behavioural change, or new variant can affect multiple implementations and product surfaces far beyond the team making it.

That does not require one central team controlling every decision. Depending on the organisation, ownership might be centralised, federated, or split between the two. What matters is where accountability sits for how the shared language evolves.

That evolution also needs to be explicit. Once tokens, contracts, and components have multiple consumers, a change may be something those consumers need to adopt. The versioning strategy will vary, but we should be able to answer what changed, which definition a consumer uses, whether the change is compatible, and what needs to migrate.

So what happens when the system can’t do something? A product team will eventually need something the design system does not provide. That is normal. A healthy system should neither force teams to fork components nor absorb every one-off need. It needs a route to recognise those gaps, then propose, review and, where appropriate, promote a reusable change – whether the result becomes a shared capability, a contextual pattern or a product-specific solution.

If an agent encounters the same situation, it shouldn’t create its own solution, it should flag it to the team as a design-system gap. The same governance process should apply whether that gap is discovered by a human or agent.

One organisation doesn’t necessarily mean one system

Large organisations often have different products, brands, and contexts. They can share foundations - common tokens, primitives, accessibility standards, and core contracts - without sharing every component or interaction pattern.

The goal is continuity without forced uniformity.

We need feedback as well as distribution

Governed definitions and contracts let us distribute knowledge across the toolchain. The other half of the problem is knowing when consumers stop following it. Without that feedback loop, small differences accumulate into drift, leaving teams with a design system that describes how the product should work rather than how it actually works.

Design-system workflows are good at pushing changes outwards; mature systems also need to detect divergence and close the loop.

Define
→
Distribute
→
Consume
→
Validate
→
Reconcile
↑
↓
└───────────────── ← ─────────────────┘
↻ repeat

Once a component has a defined contract, we can ask a series of surprisingly useful questions:

  • Does Figma expose those variants and states?
  • Does the production implementation support them?
  • Are those states represented in Storybook?
  • Does the documentation describe the same behaviour?
  • Are products consuming the approved component or recreating it locally?
  • Do the accessibility semantics match the contract?
  • If an AI agent generates a new use of that component, has it respected the same constraints?

Answering them depends on the connective tissue between systems. Naming tells us what something is, mapping tells us where its representations are, versioning tells us which definition they implement and validation tells us whether they still conform.

Some checks are already straightforward to automate. Storybook, for example, supports accessibility and visual testing against component stories.[4] The broader change is treating validation as part of the system and not just something we do only when things look wrong.

A mature design system should both distribute its decisions and detect when its consumers stop following them.

Defining what good engineering looks like

A component contract describes what a component should do; engineering rules constrain how it should be implemented in a particular environment. Those rules might cover architecture, conventions and testing, with platform-specific detail where needed.

For a product studio, this matters especially in established codebases: a design system has to coexist with the client’s engineering architecture and standards.

INTENTComponent contracts · Design tokens · Behaviour · Accessibility↓RULESArchitecture · Coding standards · Platform conventions · Testing expectations↓IMPLEMENTATIONProduction code · CI integration · Multi-model · Target platforms↓EVIDENCECompilation · Linting · Unit tests · Snapshot tests · UI testsIntegration tests · Accessibility checks · Contract validation

An extra instruction file may help a coding agent in the short term. If it becomes another stale copy of the system, however, it adds to the original fragmentation. Agent instructions, skills, and configuration should be consumers of engineering and design-system rules, rather than new sources of truth themselves.

Validation should not mean asking an LLM to check everything. Use compilers, linters and tests wherever deterministic tooling can prove the result. AI has more value when the task involves interpreting a contract, finding an appropriate component or comparing representations.

Contracts define what correct means. Rules constrain how we get there. Tests and validation provide the evidence that we did.

Greenfield ideals meet brownfield reality

It is relatively easy to draw the ideal design-system architecture on a whiteboard, and on a greenfield product you may get close to it.

Often, though, organisations are joining products that have shipped for years, with existing Figma libraries, components, tokens, documentation, and coding standards – sometimes several of each.

Existing code contains knowledge

Your existing code may contain years of decisions around accessibility, analytics, localisation, platform behaviour, and edge cases, with tests capturing behaviour that is documented nowhere else. Some of that is valuable knowledge, and some of it is historical baggage and discovery is how we tell the difference.

Before proposing a target architecture, we need to understand the current one: inspecting component APIs, dependencies, and real product usage, and talking to the people who maintain them. From there, each part of the estate can be retained, aligned, extended, consolidated or deprecated.

It also means resisting the temptation to rebuild something simply because the new version looks cleaner.

Some of the hardest design-system work involves deciding what should be left in place.

Migration is part of the architecture

Migration should not be the final box on the architecture diagram, instead the path from the existing system should influence the target architecture itself. At scale, old and new systems will often need to coexist while teams progressively converge. That’s where versioning and deprecation become especially useful. Not every difference is drift and the difference between migration and drift is whether the system knows about it.

In agency work, we are often brought in because something needs to change. That does not mean everything that came before was wrong. The job is to understand the existing context and make the next decision more coherent.

A shared contract gives platforms a common definition without requiring them to share an implementation. A Swift component and a React Native component can have different internals while following the same intended behaviour, states and accessibility requirements.

What changes when implementation gets cheap?

Sharing implementation has long been an obvious way to reduce the cost of building and maintaining sophisticated components, and that assumption has shaped component libraries and cross-platform frameworks. Generative AI has changed those economics. Code still matters, but we should question which artefacts carry the most durable value as implementation gets cheaper.

Shopify’s 2026 move from React Native back towards native Swift and Kotlin is useful evidence here. It shows how AI can change the economics of migration, rather than making a general argument for native development.[5] Its Helix tooling adds the other half of the story: small checkpoints and strict quality gates keep generated migration work shippable.[6]

For design systems, the implication is that translating a well-defined system into another implementation is now far less expensive. It is not an instruction to rewrite everything.

Perhaps the implementation isn’t the thing we need to share

Today we might share a React Native component so iOS and Android reuse the same implementation, but what if the more durable shared asset is a sufficiently precise definition of that component?

Its purpose, tokens, states, behaviour, accessibility requirements, engineering rules, acceptance criteria and tests can together describe what a correct implementation means.

Component contract+Tokens+Engineering rules+Acceptance criteria+Test suite
↓AI + automation↓Swift · Kotlin · Typescript↓Validation↓Ship

We are not completely there today, but the direction is plausible. As the cost of producing code falls, the durable value may shift from the implementation itself towards the specification, constraints and validation that tell us whether an implementation is correct.

This does not remove engineering from the process – if anything, it raises the value of good engineering. Vague contracts, incomplete tests, and contradictory rules simply let us generate inconsistent implementations faster.

Faster implementation increases the value of good constraints.

This also points to less visible uses of AI. Generating a screen makes an effective demonstration, but an agent that compares a contract with Figma and production may be more useful in everyday work. It can show teams exactly where the three have diverged.

AI as connective tissue

There is a practical problem with all of this: we have a lot of tools, each solving part of the problem well, but the answer is not necessarily another platform that replaces them all.

APIs and CI have connected these systems for years. Interfaces such as MCP now give agents another way to discover and interact with them.[2][7] Combined with coding agents, that lowers the cost of small pieces of organisation-specific tooling: token-to-Figma transforms, contract-to-Storybook validation, repo-aware implementation guidance or drift detection.

Use specialist tools for what they’re good at. Own the connective tissue that reflects how your organisation works.

That connective tissue should not become another source of truth. Agent skills, MCP servers, and instruction files should expose governed knowledge rather than redefine it.

Towards generative UI

Once a machine understands tokens, components, contracts, patterns, and engineering rules, we can ask it to do more than implement a design. We can ask it to compose an experience – and that means making decisions within the product language.

A generative system should not have to invent component meaning or guess whether a pattern is appropriate, it should understand the approved building blocks, how they can be composed, and the constraints that govern their use. And when those building blocks aren’t sufficient, it should be able to say so. That is the difference between a machine participating in the design system and one bypassing it.

Design systems are becoming infrastructure

This brings me back to Kholmatova's idea of a design system as a shared language – one that designers, engineers, platforms, products and increasingly machines all need to speak.

Tokens make decisions portable. Contracts make intent explicit. Governance and versioning control how the language evolves. Engineering rules constrain implementation. Human and machine readable views expose the same knowledge. Validation gives us evidence that everything still agrees.

That's why I think design systems are becoming infrastructure. Not an elaborate internal platform, but a dependable layer through which product decisions move between people, tools, and machines.

When that layer sits alongside the code, teams also become less dependent on any single third-party tool. Design and engineering move closer together, and roles will blur. I think that's okay. If someone without an engineering background can prototype with the same building blocks the team ships with, we can move faster without lowering the bar.

The next generation of design systems will give organisations a governed way to express how their products should work. The sooner teams invest in that, the more they'll get from it.

The real work isn't building another button. It's defining clearly enough what a button means, how it should behave, how it should be built and how we'll know the result is correct.

References

[1] Alla Kholmatova / Smashing Magazine, Design Systems. The book is framed around patterns, practices and shared language; released 2017.
[2] Figma Developer Docs, Introduction to the Figma MCP server. Structured design context and current write-to-canvas capabilities.
[3] W3C Design Tokens Community Group, Design Tokens Specification 2025.10. First stable version announced 28 October 2025; Community Group final reports. The DTCG is a W3C Community Group rather than a W3C Standards Track working group.
[4] Storybook documentation, Accessibility tests and Visual tests.
[5] Shopify Engineering, Native is now the future of mobile at Shopify, 10 September 2026.
[6] Shopify Engineering, Helix: The internal tool powering our Shopify app’s native migration, 21 September 2026.
[7] Figma Developer Docs, What the MCP sends vs. what the agent does.