LLM web-design coding tests

LLM web design test: twelve models, four identical briefs

Four briefs.
Every model got the same one.

Twelve language models received the same four web-design briefs. Same copy, same constraints, one self-contained HTML file each: no external libraries, no follow-up prompts, first output kept. The only variable is the model. Everything else — the brief, the words, the technical rules — is fixed. This page is the evidence: browse all 33 submissions, read the prompts, and compare sites side by side. For the ranked verdict, see how the models ranked.

4 briefs12 models33 pages1 HTML file each

How the test works

Each model was given the full brief as a single prompt: context, fixed copy, functional requirements, technical requirements, and an instruction to return only the HTML file. No design system, no font list, no per-model tuning. Default reasoning settings. One generation per cell.

The constraints that applied to every brief:

  • One HTML file that opens correctly from disk, CSS and JavaScript inline.
  • Vanilla HTML, CSS and JavaScript. No libraries, no CDNs, no icon packs.
  • Copy used verbatim — no invented sections, no rewritten headlines.
  • No follow-up prompts. The first complete file is the submission.

A full grid of 12 models × 4 briefs would be 48 sites. This corpus has 33. 7 models completed every brief: Claude Opus 5, Grok 4.6, GPT-5.6 Sol, Qwen3.8-Max, DeepSeek V4 Flash, GPT-5.6 Luna, Muse Spark 1.3.

The rest of the field is incomplete on purpose, not because a model refused the task. DeepSeek V4 Pro, Kimi K3 and MiniMax M3 were run on Still Running only — those are models that were available for one sitting rather than a full-time subscription. Claude Fable 5.1 and GPT-6 Astra shipped after the original four-brief run and were added on Pip & Pocket. Empty cells are briefs that were not run.

A submission is a complete HTML file. The one exception is Opus 5 on Ordinary Hours, which also shipped three hero images referenced from that file. Nothing here is a multi-turn repair, a cherry-picked rerun, or a hidden system prompt on top of the published brief.

The four briefs

Pip & Pocket site built by Claude Opus 5
01 / Early learning centre

Pip & Pocket

A 49-place childcare centre that has to feel joyful and competent at the same time.

A tired parent on their phone at 7am needs to trust this place with a two-year-old. Sites like this usually fail on one side or the other — clinical and grey, or so sugary that nothing looks credible. Holding both is the design problem.

9 submissions · 814 KBOpen brief →
Still Running site built by Grok 4.6
02 / Conference microsite

Still Running

Three days on repair, maintenance, and the work that keeps things working.

Anyone can make a design conference look designed. The question is whether you can build a visual language for bridge inspectors, night-shift mechanics and hospital biomedical engineers that is beautiful, specific, and true to the subject.

10 submissions · 600 KBOpen brief →
Three in a Row site built by Claude Opus 5
03 / Game interface

Three in a Row

Tic-tac-toe as a real product. The logic is trivial. The interface is the work.

Most game UIs are judged on the fifty seconds a player spends in menus, not the ten they spend playing. Seven screens, three palettes, three marker sets, and a computer that can actually be impossible.

7 submissions · 544 KBOpen brief →
Ordinary Hours site built by Grok 4.6
04 / Perfume house

Ordinary Hours

Three complete visual identities in one file. Same copy. Three design languages.

Anyone can produce one good-looking page. This brief asks whether you have more than one aesthetic in you — three genuinely different identities, each internally coherent, switched live from a single root attribute.

7 submissions · 435 KBOpen brief →

The models tested

Official names, labs, and release dates. Each row links to the model’s page on this site and to the vendor’s own page.

Incomplete rows: DeepSeek V4 Pro (1), Kimi K3 (1), MiniMax M3 (1), Claude Fable 5.1 (1), GPT-6 Astra (1).

How this differs from the arenas

Design Arena and WebDev Arena run blind head-to-heads. A crowd votes, identities stay hidden, and an Elo rating moves. That answers a useful question: what does average preference look like across thousands of prompts.

This is a different question. Four fixed briefs, published prompts, full source available, one judge, no voting. You can read the brief, open the file, and disagree. The arenas are a leaderboard. This is a corpus.

FAQ

Which LLM is best for frontend development?

There is no single winner that holds across every brief. GPT-5.6 Sol, Claude Opus 5 and Qwen3.8-Max each took at least one of the four tests, and the gaps were small enough that a different judge would produce a different table. The ranked write-up is in the full ranked comparison. This site is the evidence: the prompts, the constraints, and every HTML file.

How were the models prompted?

Each model received the full brief as a single prompt: context, fixed copy, functional requirements, technical requirements, and an instruction to return only the HTML file. No design system, no font list, no per-model tuning, no follow-up. Default reasoning settings. First output kept. Open any brief and use Prompt to read the exact text.

Can I see the source code for each submission?

Yes. Every submission is a self-contained HTML file. Open a cell on the board to view it live, or use the Raw link on a submission page to read the source. The original prompts sit on each brief page.

Why single-file HTML with no libraries?

A single HTML file with vanilla CSS and JavaScript isolates design and front-end craft from the rest of the stack. No component library, no build step, no follow-up to fix imports. What you see is what the model produced in one pass, which is the thing this test is actually measuring.

/ jump anywhere←→ switch modelsP read the promptC compare⇧ click a model to compare