Vibe-coding has changed a lot about what it takes to build software. Not too long ago, someone with little to no technical background would have had to sit through hours of coding tutorials, learn the basics of a programming language, and spend even longer figuring out why their code refused to work.
Now, you can describe what you want in plain English, point an AI coding agent in the right direction, and end up with something functional without understanding every line underneath it. But we’ve moved well beyond LLMs simply spitting out lines of code behind the scenes. Coding models have gotten significantly better at the visual side of development too.
Give them a screenshot of an interface, and in theory, they should be able to work backwards from the finished product like figuring out the layout, spacing, colors, typography, and components needed to recreate it. That made me wonder how far this has actually gone. So I gave Claude Code, Codex, and Google Antigravity the exact same screenshots of an app, gave them no source code or design files, and asked each one to rebuild it from scratch.
I deliberately chose an app the models were unlikely to already know
Familiarity would have ruined the test
I initially thought I’d go with an app that’s a bit more complex but also well-known. For instance, a social media app like Instagram or Snapchat, or something visually distinctive like Google Maps, where recreating the interface convincingly would actually be a challenge.
The issue is that these are all apps the models are already likely to be very familiar with. Even if I only gave them screenshots, there would always be questions of whether they were reverse-engineering the UI from what they could see or simply falling back on what they already knew the app was supposed to look like.
I thought that defeated the entire purpose of the experiment, so I intentionally went with a niche app that most people probably haven’t even heard of: Foqos. It’s an open-source focus app built around helping you block distracting apps and stay off your phone, but more importantly, for this test, it has a clean, distinctive interface without being so complex that the comparison turns into a test of backend engineering instead.
It’s also not well-known in the traditional sense, and while an LLM could technically track down its source code because the project is open source, that wasn’t something I allowed during the test. This is the prompt I used for each tool:
I used the CLI version of all three tools and set each one up in its own separate directory with the exact same set of reference images. I used Opus 5 for Claude Code, GPT-6 Sol for Codex, and Gemini 3.1 Pro for Antigravity, with higher reasoning enabled on all.
Codex was the best all-rounder
It made one hilarious mistake, though
A while ago, I wrote an article on XDA about how Codex nails almost everything except frontend work. Since then, though, OpenAI has moved on to its newer GPT-6 family, including Astra, Sol, and Luna, and the difference in frontend work is pretty noticeable. For this test, I used GPT-6 Sol, which is built specifically for more complex coding and agentic workflows.
Out of all three, Codex’s result was easily the closest to the original. Before I get into everything it did right, though, I have to point out the funniest mistake it made. Instead of calling the app Foqos, it somehow decided to rename it Fogos.
Beyond that, though, the overall proportions, spacing, card shapes, and placement of elements were impressively accurate. At a glance, it genuinely looked like the same app.
In the screenshots I provided, the Support button used a simple heart icon, and Codex reproduced it in a way that actually looked like part of the UI. Claude, on the other hand, replaced it with a literal red heart emoji, which immediately stood out against the otherwise polished interface. Antigravity was somewhere in the middle in this example.
The same applied to the other UI elements too. Codex consistently did a better job of treating the small details as actual interface components rather than approximations. Its settings icon, buttons, card outlines, and profile section all felt much closer to the reference, while Claude and Antigravity were more likely to swap in generic-looking icons or slightly reinterpret the styling. Claude, for example, simply added emojis wherever it could!
In terms of the overall fonts and typography, Codex was also the closest match. The size, weight, and hierarchy of the text felt much more faithful to the original, while Claude tended to make some elements slightly larger and Antigravity looked a little more restrained overall. None of them matched every single detail perfectly, but Codex was the only one where I could glance between the reference and the recreation without immediately noticing that something felt off.
Claude did a decent job
It got creative where I really didn’t want it to
Beyond the points I mentioned above, I think Claude did a fairly good job overall. It was actually the only one of the three that made some of the recreated controls genuinely interactive. The toggles could be switched on and off instead of just being static UI elements. Its button placement was also generally pretty close to the reference, and the overall structure of the screens made sense.
Where it fell behind Codex was mostly in the smaller visual details. The emoji-heavy icon choices were the most obvious example, but there were also a few places where the spacing, typography, and sizing felt slightly off compared to the original. Nothing was dramatically wrong, but side by side with the reference, it looked more like a very good recreation than an almost exact copy.
This was the only result where I spotted clear hallucination, though! The New Profiles section I provided only included the following sections: Name, Blocking Strategy, and Blocked Apps. I didn;t provide a scereenshot of the full page. Codex and Antigravity stuck to the options showed in the screenshot I had provided, while Claude added an Options section of its own with two toggles: Enable Live Activity and Strict Mode.
Given that those options actually exist in the real app, I got curious about whether Claude had somehow searched for Foqos despite my instructions. So, I asked it directly. Claude said it hadn’t searched the web at all and that both rows were invented based on common conventions in focus apps.
According to it, “Live Activity” sounded like a plausible feature for this kind of app, while “Strict Mode” is a common setting in app blockers. That makes the hallucination even more interesting. Claude technically broke the brief by adding UI that wasn’t visible in the screenshots, but it somehow hallucinated features that were actually there.
Antigravity had no major flaws, but it still came last
It looked good, just not close enough
Antigravity was probably the hardest one to criticize because there wasn’t really anything wrong with its result. The app looked polished, the screens were structured correctly, and nothing immediately jumped out as broken or bizarre.
The problem was that, compared with Codex and even Claude in some areas, it simply took more liberties with the design. Its cards were a little more filled out, some spacing was different, and several UI elements felt more like Antigravity’s interpretation of the screenshots than a direct recreation.
For instance, Antigravity missed some of the subtler color cues in the reference. Both Codex and Claude noticed that parts of the activity grid used slightly lighter purple shades to distinguish certain cells, while Antigravity rendered the section much more uniformly. It’s a small detail, but side by side, it made its version feel noticeably flatter and less faithful to the original.
Again, there was nothing wrong with Antigravity’s result, but there also wasn’t anything about it that made me stop and think it had truly nailed the reference. When the whole point of the experiment was to see which tool could get closest to the screenshots, those small differences added up, and that was ultimately why it came last.
Ultimately, this was a super interesting experiment because all three tools got much closer than I expected. None of them produced a perfect replica, but they were all able to infer a surprising amount from screenshots alone, including layouts, spacing, colors, component structure, and even some interactions!
