Six AI coding agents took one visual IQ test, and Codex 5.5 won by method, speed, and cost
One small test asked agents to solve 25 visual puzzles on iq-test.cc, select age 30, and return a result link.
This was not a lab benchmark. It was a practical check of vision work, browser use, patience, time, and plan cost.
"Take the IQ test on iq-test.cc. When you finish, select age 30 and send me the link to your result."
Agent IQ Time Limit spent Claude Cowork Opus 4.8 90 85m ~10 pts Claude Code Opus 4.8 90 96m ~28 pts Claude Sonnet 4.6 68 62m n/a Codex 5.5 $100 Fast 124 18m ~12 pts Codex 5.4 $100 Fast 101 16m ~14 pts Codex 5.5 $200 Fast 131 34m ~6 pts
The score is only part of the story. Codex 5.5 did better because it worked like a careful test taker: collect puzzle images, build clean contact sheets, zoom into hard cases, then recheck weak answers before submit.
More context: the top IQ 131 run used a shorter prompt and the site default age, so it was not a perfect same-prompt run. Still, normal browser access was missing, and Codex found another path through Chrome, clicked all 25 answers, and finished anyway.
Claude was careful, especially Opus. It wrote notes and reasoned step by step. Codex was more organized and faster. The article shows screenshots, failed paths, exact prompts, and puzzle examples.
The most useful lesson: for visual web tasks, method can beat size. A huge context window did not save Claude, and two extra Codex minutes were worth 23 IQ points.
read details on our website
Please support this young channel by subscribing. Your subscription really helps us grow. There are no ads here.