Most AI design workflows start with a prompt. I found that asking one model to think, remember, design, and build at the same time produces generic work — and on a regulated platform, generic is a liability.


Enterprises staff a strategist, a librarian, a systems designer, a prototyper, and a production team. This workflow gives one designer the same bench — a single tool doing each job, and judgment staying human.

 

THE PROBLEM

  • A large commercial platform with multiple user types, real money moving through it, and regulators watching every screen.
  • Eight complexity domains — real-time systems, compliance, financial logic, authentication, identity, API integrations, media, permissions — with [ # ] backlog items rated 4/5+ for difficulty and [ # ] third-party APIs. Each domain has its own failure modes.
  • When AI lacks context it reaches for the generic answer. A generic answer to a permissions problem or a piece of financial logic doesn’t survive its first compliance review.
  • Context drift is AI’s most expensive weakness at this scale. A conversation ends and walks off with its reasoning — a month later nobody can reconstruct why the permissions model works the way it does.

 

GOALS

  • Run a stack of AI tools as a multidisciplinary design team — one responsibility per tool, hard lines between them, judgment at the center of the process.
  • Get through more good decisions per week without dropping the bar, and hand engineering something unambiguous instead of a folder of static screens.

 

MY ROLE

  • UI/UX Product leader, Architecture designer, Lead prototyper, and AI connection specialist.
PROJECT RESULTS

Impact and Outcomes

Stakeholders reviewed a polished, coherent, high-fidelity artifact and reacted to the product instead of imagining it — the review bar this workflow lives or dies by.

One job per tool: strategy, memory, system definition, functional prototyping, and high-fidelity design each lived in a separate, specialized layer — so no platform was ever asked to think, remember, design, and build at once, and nothing drifted off-system.

Executable specs, not mockups: working HTML, CSS, and JavaScript prototypes proved the logic, state, and validation held up long before development started — not just that the layout looked right.

Memory that outlasts the conversation: every significant product decision was logged as it was made, in one fixed format — by the end the decision log was worth as much as the backlog sitting next to it.

Metric OneMain-style (lending) TIAA-style (tax/retirement)
User stories ~180 ~140
Business rules ~200 ~120
User types 11 9
Risk domains 9 8
Complexity domains 10 8
Third-party APIs 12 8
Items rated 4/5+ difficulty ~150 ~100
Decisions logged ~150 ~110
Backlog items resolved before engineering ~265 (~70%) ~180 (~70%)
Stakeholder review cycles before production code 3 3
To comply with my non-disclosure agreement, I have omitted and obfuscated confidential information in this case study. The information in this case study is my own and does not necessarily reflect the views of the clients this workflow has been built for.
1 of 7 7%
THE APPROACH

Dividing the Work, Don't Prompt Yet

I didn’t ask one AI to do everything. I gave each tool a single job, drew hard lines between them, and kept judgment at the center where it belongs. Over the past several months I led the end-to-end design of a large commercial platform.

I can’t tell you what it was. I can tell you how it got built. The move that mattered: I stopped treating AI as a faster design tool and started treating each tool as a specialist collaborator with one clear responsibility — then refused to let any of them freelance outside it. This was never about automating myself out of the work. It was about getting through more good decisions per week without dropping the bar.

Every call still ran through me. The tools just closed the gap between making a decision and seeing it come to life.

2 of 7 28%
THE STATUS QUO

Screens are Easy. Managing Complexity is the Job.

The “before” state isn’t a tool — it’s a habit. Open one AI, ask it to think, remember, design, and build in the same thread, and watch it default to the generic answer the moment the context runs thin.

On a platform with eight complexity domains, that default is the risk. Old: one chat window, one model, one context that quietly drifts. New: five specialized layers, each with a single responsibility and a clean handoff to the next — strategy, memory, system, prototype, high-fidelity design. Real-time systems. Compliance. Financial logic. Authentication. Identity. API integrations. Media. Permissions.

Here’s the thing about AI and context: when it doesn’t have enough, it reaches for generic. On a platform like this, generic isn’t good enough. It’s a liability.

3 of 7 42%
MAKING IT REAL

Dividing the Work

Every tool should do one thing exceptionally well. I never asked a single platform to think, remember, design, and build at the same time. Each layer of the workflow had its own specialist, and I policed the borders between them.

Layer Tool Its one job What it never did
Strategic thinking Claude Feature definition, information architecture, user journeys, interaction models, and the trade-offs between them — in conversation, before a single screen existed. Produce UI. This stage generated zero pixels on purpose.
Persistent memory Notion Backlog items, decision logs, feature specs, and business rules captured as structured knowledge while the decisions were still warm. Get opened by me. Claude pulled from it and reasoned over it in conversation.
System definition Figma Source of truth for the visual language — typography, spacing, interaction patterns, components, the design system itself — locked before anything downstream touched it. Exploration. By the time something landed in Figma, the thinking was done.
Functional prototyping Claude Translate the locked system into working HTML, CSS, and JavaScript to test real interactions, state logic, and validation long before development started. Stand in for a mockup. Each one was closer to an executable spec.
High-fidelity design and handoff Claude Design Compose the resolved system and validated interactions into polished screens and flows on a canvas driven by chat — the artifact stakeholders reviewed and developers built from. Deploy or replace engineering. It’s a design surface, not a pipeline. Production stayed with engineering.

Strategic thinking — Claude. Claude was my product strategist. Before a single screen existed, we talked through feature definition, information architecture, user journeys, interaction models, and the trade-offs between all of them. This stage produced zero UI, and that was the point. I was building understanding, not pixels.

Persistent memory — Notion. Notion held the project’s long-term memory. The part that surprised me: I almost never opened the backlog myself. Claude pulled from it and reasoned over it in conversation, so the backlog stopped feeling like software I had to drive and started feeling like context that was simply there whenever we worked.

System definition — Figma. Once the direction was set, Figma became the source of truth for the visual language. I kept exploration out of Figma on purpose. Figma’s job was to pin the system down with precision.

Functional prototyping — Claude. Claude translated the locked design system into working code. “Mockups” is the wrong word for what came out of this. Each one was closer to an executable spec — proof that the logic held up, not just that the layout looked right.

High-fidelity design and handoff — Claude Design. Claude Design came in last, once the hard calls were behind us: strategy set, interaction model agreed, design system locked. Its job wasn’t to invent anything. Because the thinking was already done and the system was the input, this stage moved fast and stayed on-system by default. Nothing drifted. What came out was the artifact stakeholders actually reviewed — close enough to the real thing that people reacted to the product instead of imagining it — and a far richer reference for handoff than static screens ever could be. Engineering still owns production. This just made sure they started from something unambiguous.

4 of 7 57%
INSTITUTIONAL MEMORY

A Decision Log Worth as Much as the Backlog

Context drift is AI’s most expensive weakness on a project this size. A conversation ends and walks off with its reasoning, and a month later nobody can reconstruct why the permissions model works the way it does. So I logged every significant product decision as I made it, in one fixed format:

  1. Decision
  2. Context
  3. Alternatives considered
  4. Trade-offs
  5. Consequences
  6. Related features
  7. Dependencies

By the end, the decision log was worth as much as the backlog sitting next to it — and because Claude reasoned over both in conversation, neither one became software I had to drive.

5 of 7 71%
WHERE THE DESIGNER SITS

The Designer Moves Upstream

This project hardened a belief I already held. As AI gets better at generating interfaces and turning designs into code, the difference gets made by the quality of the thinking that happens before any of that generation starts.

The designer’s job follows that shift — less time producing screens, more time defining systems, framing problems, managing ambiguity, and holding a coherent product direction steady while everything around the work speeds up.

Execution keeps getting cheaper. Judgment hasn’t moved an inch.

6 of 7 85%
WHAT I KEPT

The Product Belongs to the Client.
The Workflow Came Home With Me.

The workflow keeps strategic thinking, memory, system definition, functional prototyping, and high-fidelity design in separate, specialized layers. That’s what lets AI accelerate execution without the product paying for it in quality or design intent.

The tools will churn. Half this stack could look different in a year. The division of labor underneath them is the part I expect to keep.

7 of 7 100%
AFTER THE TIME SAVINGS

So What Do You Do With the Time You Get Back?

This is the question most teams skip, and it’s the one that decides whether any of this was worth it. The wrong move is to point all the saved time at shipping sooner. Speed on its own just gets you to the same place earlier — mistakes and all. When a prototype costs an afternoon instead of a sprint, “let’s just try it and see” becomes a reasonable thing to say out loud, and you learn things you’d otherwise have shipped right past.

Where I’d spend the surplus, in order:

  • Test what there was never room to test. The assumptions that got cut for time on every previous project are now an afternoon each.
  • Chase the small, messy problems. The customer problems that never made the roadmap because they couldn’t justify the old build cost — now they can.
  • Ship sooner — where the case is clear. Only where the risk is already understood. The real return on cheaper execution isn’t more output. It’s better questions.
  • Quality has one owner. None of it changes who’s accountable. I am. The tools don’t carry the risk, they don’t answer for a confusing flow or a compliance gap, and they can’t be held accountable to a user. Faster generation raises the bar on judgment — it doesn’t retire the person applying it.
PROJECT RESULTS

Impact and Outcomes

Stakeholders reviewed a polished, coherent, high-fidelity artifact and reacted to the product instead of imagining it — the review bar this workflow lives or dies by.

One job per tool: strategy, memory, system definition, functional prototyping, and high-fidelity design each lived in a separate, specialized layer — so no platform was ever asked to think, remember, design, and build at once, and nothing drifted off-system.

Executable specs, not mockups: working HTML, CSS, and JavaScript prototypes proved the logic, state, and validation held up long before development started — not just that the layout looked right.

Memory that outlasts the conversation: every significant product decision was logged as it was made, in one fixed format — by the end the decision log was worth as much as the backlog sitting next to it.

Metric OneMain-style (lending) TIAA-style (tax/retirement)
User stories ~180 ~140
Business rules ~200 ~120
User types 11 9
Risk domains 9 8
Complexity domains 10 8
Third-party APIs 12 8
Items rated 4/5+ difficulty ~150 ~100
Decisions logged ~150 ~110
Backlog items resolved before engineering ~265 (~70%) ~180 (~70%)
Stakeholder review cycles before production code 3 3
To comply with my non-disclosure agreement, I have omitted and obfuscated confidential information in this case study. The information in this case study is my own and does not necessarily reflect the views of the clients this workflow has been built for.