THE MANUAL · AI-NATIVE DESIGN · BRIEF → SHIPPED
Design with AI,
end to end.
Start with one line. Ship something real.
Ten steps, written for someone who has never designed with AI before. The full landscape of tools, each with a when-to-pick-it line. Working templates, ready to copy. The failure modes of every step, and the test that tells you when to move on. The philosophy behind it lives in the companion essay.
STEP 01 / 10
Brief
Decide what done looks like.
Everything downstream inherits the brief. A vague brief gives you vague screens, faster - and AI will happily design the wrong product beautifully. The brief is the one artifact that is entirely yours, and it is short: who it is for, what they are trying to do, when it fails today, what done looks like. Do not open a design tool until it exists.
THE TOOLS · AND WHEN TO PICK WHICH
THE TEMPLATES
I want to design [the thing]. Interview me one question at a time until you can write a one-paragraph brief that names: who it is for, what they are trying to do, when it fails them today, and what "done" looks like. Then write the brief.
For [who - one kind of person, not "users"]. They are trying to [the job, in their words]. Today it fails when [the specific moment it breaks]. Done looks like [an outcome a stranger could verify]. Not in scope: [what you are NOT designing - one line that saves you later].
WHERE THIS STEP GOES WRONG
- Writing the brief as a feature list. Features are answers; the brief is the question.
- Skipping the not-in-scope line. Scope creep starts here, not in the build.
- Letting the AI write the brief without the interview. It fills your gaps with averages.
What good looks like: a stranger reads it and describes the same product back to you.
DONE WHENYou can read the brief to a stranger and they can describe the product back to you.
STEP 02 / 10
Eval
Write the test before the answer.
ZURB’s line, and the best idea in design right now: the eval is the new spec. When a prototype costs an afternoon, the scarce work is judging which one ships. Write the criteria now, while you are still impartial - after you see directions, you will write criteria that fit your favorite. Three to five criteria, each one observable, each one scored 0 or 1.
THE TOOLS · AND WHEN TO PICK WHICH
THE TEMPLATES
A direction passes when: 1. A first-time user can [core job from the brief] in under [N] steps. 2. The first screen answers "what is this, and is it for me" in 5 seconds. 3. [One criterion of your own - make it observable.] Score every direction 0 or 1 per criterion. Below full marks: iterate, don’t ship.
Scorecard - one row per direction, one column per criterion. Direction | C1 | C2 | C3 | Verdict A | | | | B | | | | C | | | | Rule: the verdict is written from the scores, not from memory of which one looked nicest.
WHERE THIS STEP GOES WRONG
- Criteria you cannot observe. "Feels premium" is not a criterion; "loads in under 3 seconds on a phone" is.
- Writing evals after seeing the directions. That is rationalizing, not evaluating.
- More than five criteria. Past five, none of them bite.
What good looks like: two people score the same screen the same way without talking.
DONE WHENThe criteria are specific enough that two people would score the same screen the same way.
STEP 03 / 10
Research
Collect proof before you invent.
This is not academic research. You are loading two contexts with real decisions: yours, and the machine’s. Two kinds of proof matter - what users actually do (evidence) and what real products already do (references). An hour here saves days of generating plausible nonsense.
THE TOOLS · AND WHEN TO PICK WHICH
THE TEMPLATES
Reference: [app + screen] Problem it solves: [one line] What I am taking: [the decision, not the pixels] What I am leaving: [what does not fit my brief]
Three questions this design must answer: 1. [question] - what I believe: [guess] - how I will check: [source or test] 2. [question] - what I believe: [guess] - how I will check: [source or test] 3. [question] - what I believe: [guess] - how I will check: [source or test]
WHERE THIS STEP GOES WRONG
- Trusting AI summaries without opening the sources. Fluency is not evidence.
- Collecting forty references and reading none. Ten studied beats forty saved.
- Copying patterns from products with different users and different stakes.
What good looks like: every reference has one line on the decision you are taking from it.
DONE WHENFive to ten annotated references, and your three beliefs with a way to check each.
STEP 04 / 10
Structure
Name things before you draw them.
Information architecture: what exists, what it is called, and how it connects. AI generates screens happily with no structure at all - you get beautiful orphans. Decide the objects, the navigation, and the names first. Names matter twice now: users read them, and the agents you prompt inherit them.
THE TOOLS · AND WHEN TO PICK WHICH
THE TEMPLATES
Here is my brief: [paste it]. List every screen and state this product needs. For each: its name, its one job, and what it links to. Flag anything the brief does not cover. Keep it under [N] screens - cut before you add.
- Every name is a word the user already uses, not a word from your org chart - Top level has 7 or fewer things - Every screen has exactly one job - You can draw the whole product on one sticky note
WHERE THIS STEP GOES WRONG
- Designing screens before listing them. The list is the design.
- Naming from your org chart instead of the user’s vocabulary.
- Eleven top-level sections. If everything is important, nothing is.
What good looks like: the whole product fits on one sticky note and a stranger can navigate it.
DONE WHENA one-page sitemap where every screen has one job and a user-word name.
STEP 05 / 10
Directions
Generate five, not one.
Volume is the point. The same prompt into three tools gives you three schools of thought, and the differences teach you what you actually want. Keep the prompt identical across tools - you are comparing directions, not tools. Work screen by screen: asking for the whole app at once gets you an average of everything. This is the whole prompt-to-UI landscape - pick three, not eight.
THE TOOLS · AND WHEN TO PICK WHICH
THE TEMPLATES
[Paste the brief from step 1] [Paste the eval criteria from step 2] Generate one direction for the first screen. Prioritize criterion 1 over polish. No lorem ipsum - write the real microcopy. Show me the state the screen is in when it is doing its job, not an empty state.
One line per direction, written BEFORE scoring: A: [what this direction believes about the user] B: [same] C: [same] Then the scorecard from step 2 decides. If two tie, the one that is easier to explain wins.
WHERE THIS STEP GOES WRONG
- Falling for the first good-looking output. Pretty is not passing.
- Tool tourism - eight builders, eight different prompts, nothing comparable. Pick three, same prompt.
- Asking for the whole app in one prompt. Screen by screen, always.
What good looks like: five directions different enough that scoring them teaches you something.
DONE WHENFive or more directions you could defend or kill with the eval from step 2.
STEP 06 / 10
Visual design
Pick one. Make it yours.
Score the directions, pick the winner, and steal the best idea from each loser into it. Then refine - this is where taste enters, and the step you do not delegate. Type scale, spacing rule, color with a reason, assets that belong to this product. AI is excellent at variations and mediocre at restraint; restraint is your job.
THE TOOLS · AND WHEN TO PICK WHICH
THE TEMPLATES
Keep the layout. Fix exactly these things: [specific, numbered list]. The type scale is [your scale]. Spacing follows [your rule]. Do not redesign - show me the before and after of each fix.
Type: [3 sizes max - display, body, caption] in [one typeface] Spacing: [one base unit - every gap is a multiple of it] Color: [one ink, one paper, one accent - accent means one thing] Icons: [one family - name it] Rule: if an element needs a new token, the element is probably wrong.
Generate an image for [exact slot in the design]. Style: [three adjectives that match the brief]. Palette: [your tokens]. It must read at [size]. No generic stock-photo people, no gradients I did not ask for. Give me four variants - I will art-direct from there.
WHERE THIS STEP GOES WRONG
- Letting the tool redesign when you asked for a fix. Name the fix; keep the layout.
- AI imagery that could belong to any product. If the asset has no reason, the answer is no asset.
- Mixing icon families and font voices. One family each - mixing reads as accident.
What good looks like: every element traces to the brief or the eval, and nothing is decoration.
DONE WHENEvery element on the screen traces to the brief or the eval. Anything that cannot is deleted.
STEP 07 / 10
Prototype
Make it run.
A prototype that runs beats a picture of one. Real data, real states, real transitions - this is where AI-native compresses weeks into an afternoon. Stay where the direction already runs, or hand the refined design to a coding agent. Either way: match the design exactly before improving anything, and make the handoff explicit - handoff is where design quality goes to die.
THE TOOLS · AND WHEN TO PICK WHICH
THE TEMPLATES
Here is the design: [link or screenshot, plus your type and spacing tokens]. Build it as [your stack]. Match the design exactly before you improve anything. When you hit a judgment call, ask me instead of deciding silently.
- The core job from the brief works end to end on your machine - Real data, not placeholder - even if you typed it yourself - Loading, empty, and error states exist and say something useful - It runs on your phone, not just your laptop
- Tokens named the same way in design and code - Every interactive state specced: hover, loading, empty, error - The builder saw the prototype run, not just pictures of it - One named owner for design decisions after handoff
WHERE THIS STEP GOES WRONG
- Improving the design while building. Match first, improve second - one variable at a time.
- Letting the agent decide silently. Make it ask; silent decisions compound.
- Demo data that flatters the design. Real mess is the test.
What good looks like: the core job works with real data while you watch, on a phone.
DONE WHENThe core job from the brief works end to end on your machine.
STEP 08 / 10
Test
One stranger beats ten iterations.
You do not need a lab. You need one person who did not make the thing, three tasks, and the discipline to stay quiet. Two tests carry most of the weight: the five-second test (what is this, is it for me) and the task test (can they do the core job unaided). Their struggle is the data - helping them erases it.
THE TOOLS · AND WHEN TO PICK WHICH
THE TEMPLATE
Setup: "I am testing the product, not you. Think out loud. I cannot help." 1. Five-second test: show the first screen for 5 seconds. Ask: what is this, and is it for you? 2. Task test: [core job from the brief], then [second job], then [edge case]. Watch. Say nothing. 3. Close: what would you tell a friend this does? Write down every hesitation with a timestamp. Hesitations are the next brief.
WHERE THIS STEP GOES WRONG
- Helping when they struggle. The struggle IS the data.
- Testing on friends who already know the product. Politeness is noise.
- Asking "do you like it". Watch what they do; ignore what they say they would do.
What good looks like: a stranger completes the core job unaided and describes it back correctly.
DONE WHENOne stranger used it and you have a written list of where they hesitated.
STEP 09 / 10
Ship
A URL beats a folder.
Ship the same day. A live URL changes what the work is - people can use it, share it, and break it, and every one of those teaches you. Done is not "looks right on your machine"; done is a stranger using it without you in the room. Ship means three things: it is up, it is measured, and everyone can use it - including the people your design usually forgets.
THE TOOLS · AND WHEN TO PICK WHICH
THE TEMPLATE
- Loads on a phone in under 3 seconds - One person outside the project completes the core job without help - Every number and claim on the page is true and checkable - An automated accessibility pass shows no errors on the core flow - Analytics answer one question: did they do the core job - There is a way for a user to reach you - The URL has been shared with someone who did not make it
WHERE THIS STEP GOES WRONG
- Waiting for perfect. Perfect is a loop count, not a launch date.
- Checking accessibility after launch instead of before. The checklist is cheaper than the fix.
- Analytics with no question attached. Pick the one number that proves the core job.
What good looks like: the URL works for a stranger on their phone on their network - including a stranger using a screen reader.
DONE WHENThe URL is live, measured, accessible - and someone who did not make it has used it.
STEP 10 / 10
Iterate
The loop is the discipline.
The last line of your first run - where the stranger hesitated - is the first line of your next brief. That loop is the whole discipline. Change one thing per loop where you can, keep the eval criteria from step 2 in force, and let evidence, not enthusiasm, set the order.
THE TOOLS · AND WHEN TO PICK WHICH
THE TEMPLATE
Loop [N] retro: What they did: [observed, not interpreted] What I expected: [be honest] What changes: [one thing] What stays: [what the eval still protects] Next brief’s first line: [the sharpest hesitation, in their words]
WHERE THIS STEP GOES WRONG
- Collecting feedback you never read. Hoarded signal is zero signal.
- Changing ten things at once. Then no result means anything.
- Dropping the eval criteria once the thing exists. They age; refresh them, do not abandon them.
What good looks like: every loop starts from evidence and ends with one deliberate change.
DONE WHENThe second brief exists, and its first line came from a user, not from you.
AFTER STEP 10 · THE LOOP
Then do it again.
Every loop starts from evidence and ends with one deliberate change. Everything else on this site feeds it: 78 live patterns to try in your browser, the readings in order, the rankings for choosing tools, the rules I do not delegate, and the tools I actually watch.
Back to the wall ↗