How the tests are arranged
The author introduces GPT-6 Astra through a series of his own tests, then turns to published benchmark results and access conditions. He does most of the work in the ChatGPT desktop application, which he describes as the former Codex: he explains that agents there can work on several projects simultaneously and use persistent local files. Instead of familiar demonstrations, he chooses tasks involving physics, animation, operating interfaces, and producing finished materials. He also repeatedly assigns a separate agent to evaluate the result and return a list of shortcomings to the builder for another attempt.
A simulation of a bullet and water
The first test is a simulation, written from scratch, of a bullet piercing a water-filled balloon. After the balloon bursts, the water should change shape and fall to the ground; the result should look as realistic and physically accurate as possible. The author prohibits external web libraries, including the 3D library named in his prompt, and requests the ability to pause at any moment, choose any viewing angle, and adjust sliders for bullet speed, balloon size, lighting, and gravity. A separate critic must take screenshots from different sides and rate the physics and realism on a scale from 0 to 10. A passing result requires at least 8 points, with a maximum of three attempts; the author selects ultra mode for the run.

Work on the simulation takes 32 minutes. The critic gives the first version 4.7 points, the second 6, and the third 6.5. None reaches the required 8, but the process stops after the third round, as instructed. In the resulting application, the author changes the balloon diameter, bullet speed, impact height, gravity, surface tension, and spray density. He moves the timeline, demonstrates the bullet passing through the water and the water falling, shows side, top, and bottom views, and changes the light intensity and direction. He considers the result a reasonably decent-looking simulation written without external libraries. He cannot show the number of tokens used because the application does not display session statistics. With several tasks running simultaneously, he estimates that this particular job consumed 3–4% of his weekly allowance.
A game in Unreal Engine
The next task is a procedural Unreal Engine game with a character controlled from a third-person viewpoint. The author considers demonstrations of shooters and racing games less revealing because, in his view, they involve relatively little complex animation and articulated movement. Here, the model must take a character from Sketchfab, find animations on Mixamo, and apply them to the character: the heroine should sprint and leap across rooftops like a ninja. For assets requiring authentication, it may use the Playwright extension and the author’s existing Chrome session, where he is already logged in. The environment is specified as an ancient Chinese imperial complex, with a previously generated image as a visual reference. All assets must be created through Blender MCP and placed procedurally in Unreal Engine, aiming for a detailed, realistic, and grandiose scene that makes an impression comparable to an AAA game.
The game’s critic must independently capture screenshots from several viewpoints and at different zoom levels, score the design and appearance from 0 to 10, and provide a ranked list of fixes. Success requires a score above 8.5 with no errors, and the limit is four rounds. The first stage takes an hour and 26 minutes: the agent downloads the character, adds running and airborne poses, connects movement to Unreal, and creates buildings. The critic notes a rather empty scene, excessively metallic roofs, repetitive layouts, and crude distant mountains. After revision, the second version receives 5.8 points; the final round reaches 6.3. The threshold is again missed, and the author finds the character’s movements particularly unsuccessful.
The author continues with additional instructions. He asks for more suitable Mixamo animations, a forward lean during fast running and jumping, further critic reviews—again up to four rounds—and expressive camera movement. The excessively high single jump should be reduced, while chained jumps in mid-air should become possible; the running and jumping animations should play twice as fast. He also requests camera shake, motion blur, wind or mist, more dramatic and realistic lighting, varied buildings, and an infinite procedural palace. After approximately four hours, he demonstrates a game in which the player can walk, run, cross rooftops with chained jumps, and collect floating icons added by the agent. He likes its appearance, although some building floors and lanterns remain disconnected. He also shows the wireframe view and a lighting-only view. He estimates consumption for roughly four hours at about 10% of his weekly allowance, again without available session statistics.
Drawing and performing music through an interface
In the computer-use tests, the author first mentions completing tax forms and editing Microsoft Office files, then demonstrates drawing in Photopea. The agent must open the website, create a blank canvas, and reproduce an attached picture in as much detail as possible. The accelerated recording shows individual brushstrokes; the author emphasizes that the agent uses the cursor rather than constructing the image programmatically. The process takes approximately 20 minutes. He then asks the agent to draw frame-by-frame sprite animations in another web editor, using a character image as reference: running, jumping, and a sword slash, each comprising four frames and saved as a separate GIF. In roughly 40 minutes, the agent draws the movements sequentially, frame by frame, and delivers three animations.
The author considers playing a virtual piano a more difficult interface test. The prompt asks the agent to enable the maximum number of visible keys, compose a fast-paced solo in the style of Chopin lasting about a minute, and perform it by pressing the online keyboard’s keys. Besides composition, the task requires maintaining the timing between successive key presses rather than simply playing prepared audio in a music application. The agent spends several minutes figuring out the keyboard and recording its performance; the author thinks there is a short rehearsal first. The finished piece then plays. He says it sounds like a legitimate classical piano solo and notes a total working time of about seven minutes.
A composition in Waveform
The author commissions a complete composition in Waveform DAW: europop EDM using independently selected instruments, available VST plugins and samples, variations, risers, impactful drops, panning, effects, and automation. The request also calls for professional mixing and mastering. The first render arrives after 15 minutes, but the author finds it basic and amateurish, with an unnatural-sounding drop impact. In a second prompt, he allows free samples from samplefocus.com and asks for more interesting sonic details and a catchier melody. Approximately 20 minutes later, the revised version plays. He likes it substantially more than results previously obtained from Claude and GLM, while explicitly acknowledging the additional instruction and permission to find new samples. Even the improved track, in his assessment, sounds somewhat muddy, is not mixed well enough, and falls short of professional human mastering.
Reconstructing a building in Blender
To test 3D reconstruction, the author supplies an entire Airbnb listing page for expensive accommodation in Japan. He points out a slanted wall beside the pool, an unusual wooden object, the bathroom, the building’s exterior, the living room, a long main pool and a smaller circular pool, a sofa, and a table. Through Blender MCP, the agent must reconstruct the building and its interior, program a camera path for a virtual tour, and render the scene. The work takes an hour and 12 minutes, with estimated consumption of 5–10% of the weekly allowance. Thousands of elements are created in the scene. The author shows the bathroom, pool area, and living room in wireframe and standard views, then displays the complete render beside photographs of the property. He considers the result quite decent, although it does not fully match the original.
A commercial and a mathematical video
Another task tests the model’s role as a director controlling other generators. The author says that such models do not create images and videos themselves but can direct specialized tools. He provides an Amazon product page and requests a 30-second commercial, using the photographs and specifications as reference. Content must be generated through Higgsfield MCP; the agent may choose suitable image and video generators, create several clips, and combine them. A finished commercial appears after approximately 20 minutes. The author rates it substantially above the result Claude Fable 5.1 produced from the same prompt.
The agent then creates an approximately one-minute mathematical explanation, from scratch, of how Earth’s circumference was first determined. The requirements are minimalist white graphics on a black background, diagrams, accessible presentation, and sufficiently detailed mathematics. Tools may be chosen freely; for the voiceover, the author specifies Gemini TTS and supplies documentation. He considers the initial version, produced after 14 minutes, acceptable but asks for improvement through a critic: at least 8 points, with a maximum of three rounds. Later, he separately requests a more coherent script suitable for explaining the topic to high-school students. The final video presents Eratosthenes’ method around 240 BC: at noon on the summer solstice, a stick in Syene casts almost no shadow, whereas a stick in Alexandria does. Dividing shadow length by stick height gives the tangent of the angle; the modern explanation proposes finding the angle with inverse tangent. The angle of approximately 7.2° becomes the angle between Earth’s radii under nearly parallel sunlight. It is one-fiftieth of a full 360° circle; multiplying a distance of approximately 800 kilometers by 50 gives about 40,000 kilometers. The video itself acknowledges that the measurements were approximate.
Finding a frog and identifying tumors
In the online interface, the author switches to the Work tab and tests finding a hidden frog in a photograph. He asks the model to identify and circle any animals and selects max mode. After 6 minutes and 50 seconds, the model does not find the frog: it says it cannot confidently identify an animal and suggests a rattlesnake, which the author calls incorrect. He mentions users’ reports of successful attempts but says he could not reproduce them over several runs. In another answer, the model incorrectly indicates a small toad; on another photograph from Frogbench, it also circles the wrong place. Based on his attempts, the author concludes that the frog test has not been passed.

Next, he uploads six brain images, each containing a different tumor type according to his account, and asks the model to name any tumors present. The task takes two and a half minutes. The model correctly identifies the upper-left image as a meningioma. It finds no definite tumor in the upper-middle image and again names a meningioma in the upper-right one; it treats the lower-left and sixth images as having no tumor, and suggests a calcified meningioma or the option written as “cranio” for the lower-middle image. The author considers all five of these answers incorrect. The result is one correct answer out of six, matching Kimi and GLM and, in his comparison, representing the highest score for this prompt.
Research and new ideas
The research task asks for an analysis of amyloid-beta and tau propagation in Alzheimer’s disease, a comparison of therapies targeting each protein, and critical appraisal of recent phase 3 trial outcomes, with tables and visualizations. The author believes there is no single correct answer here, and that most frontier models apart from Claude perform well, differing primarily in presentation. In Astra’s response, he shows a table of mechanisms, source citations, a flowchart, a table of phase 3 trials, a programmatically created figure, and another information-dense table. His positive assessment concerns concision and information density.
The final original prompt requests three ideas that do not yet exist and would have the greatest impact on the quality of life of animals in factory farms, using current technology and AI. The constraints require meaningful improvement, application across enormous numbers of animals, low labor and operating costs, and obvious appeal to operators. After 8 minutes and 35 seconds, the model proposes creating and protecting resting zones for chickens, adapting enrichment for pigs to demand and competition, and preserving compatible pig groups during routine moves. For the first idea, it lists less interrupted rest, greater choice over surroundings, potential feed savings, and fewer leg problems. The author explicitly acknowledges that evaluating these proposals is subjective.
Reported benchmarks and gaming achievements
Turning to benchmarks, the author initially identifies the results as developer-reported. He gives approximately 98% on Frontier Math and describes Astra as leading Terminal Bench 4, Automation Bench, Agents last exam, and OS World. He associates these tests with difficult mathematics, agentic terminal work, multi-step business processes, professional tasks, and operating computer applications, respectively. He relays the claim of the world’s best computer-use model as a claim made by others. He also shows leadership in reconstructing 3D objects from different views and improvement over the previous GPT on String Quartets, which converts audio into musical notation.
Among the gaming achievements he cites, the author calls Astra the only general agent at that point to have completed Portal. He explains the significance in terms of spatial reasoning, physics understanding, multi-step planning, and precisely timed mouse and keyboard actions. For Pokémon Fire Red, he gives a completion time of 18 hours and 12 minutes, substantially shorter than those of other models. He connects completion to long-term planning, memory, vision, and thousands of coherent decisions. He then mentions a Pokémon Crystal stream on Twitch and predicts a future record; the presentation contains no completed Crystal result.
For ARC AGI-3, the author gives a result above 60%, compared with below 10% for other frontier models, and nearly 100% with an additional harness. He describes the test as learning the rules of an unfamiliar environment and explains the difficulty for models through weights that remain fixed after training. On Voxel Bench, for 3D voxel animations, Astra ranks first in the displayed table and leads Opus 5 by almost 400 points, while the other models differ by a few dozen. On one of the cited Frontier Math tests, it is the only model scoring above 0% on problems the author describes as previously unsolved. On iBench, which tests vision through mazes, spotting differences, graphs, and patterns, he gives 95%, compared with below 50% for others. He also notes first place in New York Times Connections puzzles.
Not every table presented puts Astra first. In the Artificial Analysis Intelligence Index, the Max version ranks second behind Claude Fable 5.1; the author nevertheless considers its cost per task reasonable and Claude excessively expensive. The displayed Omniscience measure shows fewer hallucinations for GPT-6 models than for Opus 5, Fable models, and the earlier GPT-5.6. On LiveBench, Astra ranks third, below Fable. In Arena, it has not yet been added to some categories, but the author shows first place with a large lead in WebDev.
Access and use of the weekly allowance
At the end, the author says access should have reached all paid ChatGPT plans, including Plus and Pro, by the time of recording. He uses Pro 5X himself: he cannot select GPT-6 in the ordinary online chat, but it is available to him in the Work tab and desktop application. He considers his plan’s weekly allowance reasonable and estimates that a typical hour-long coding task uses approximately 5–10% of it; for Plus, he anticipates substantially faster consumption. His closing assertion that it is the most capable and powerful available model remains his own assessment, based on the demonstrated tests and comparisons he presents.



