Google DeepMind has introduced a database predicting the effects of nine billion single-letter changes in human DNA. OpenAI has announced a mathematical result pursued by ten thousand concurrent agents. Meanwhile, a new benchmark using real company code left leading models below a 40% task-resolution rate. September’s releases bring large-scale scientific computation alongside specific actions: changing a word in an existing recording, transferring human movement to a robot, or loading only the parts of a language model currently needed.
Marigold V2: depth and surface structure from an ordinary image
Marigold V2 is presented as a model that extracts three-dimensional depth and structure from an ordinary image. Its input is a picture; its outputs include a depth map, surface normals, albedo, and other maps. The presenter explains normals as the orientation of a surface within an image and albedo as its raw color.
The presenter associates the new version’s main difference with pixel-level operation. In the comparisons, he emphasizes detail, sharpness, and accuracy as well as overall object shape. In several displayed comparisons with other models, he describes Marigold as more detailed and higher resolution. Separate examples compare normal estimation, which he considers sharper and more faithful to the image. He also says Marigold V2 obtains the best average scores among the comparable models in the displayed benchmarks.
Code has been released, with download and local execution instructions in the project repository. The displayed settings require about 17 GB of VRAM for 1024 × 1024 images and about 29 GB for 2048 × 2048. Those figures refer to the two listed processing settings. The more detailed map requires more VRAM.
UniMate: motion for very different skeletons
UniMate addresses another three-dimensional task: animating characters whose skeletons have different structures. The presenter contrasts it with models specialized for humans that often struggle with birds, snakes, and other unusual objects. The demonstration uses one model for flowers, Garfield, a satellite, and a dragon-like creature. The variety matters: these are not merely human characters of different heights but substantially different forms. In the presenter’s assessment, the model handles all the displayed animations well.
Generation requires a rigged 3D model and a text instruction. The examples are a dog walking forward and a snake slithering forward. According to the presenter, UniMate generates the motion automatically without additional retraining for each such object. He calls it one of the best choices for unusual characters and objects; that is his assessment of the displayed examples. Training code has been published alongside local execution code, which is why he describes the project as fully open source.
AlphaGenome Atlas: predictions for billions of DNA variants
According to the presenter, Google DeepMind has released AlphaGenome Atlas, a large AI-generated map of human DNA. To explain the scale, he says roughly 2% of the genome, mostly associated with protein coding, is better understood, while the remaining 98% is much more mysterious. The Atlas task is described as predicting the effects of every possible single-letter mutation, amounting to approximately nine billion variants. He stresses that experimentally testing all of them would be practically impossible: these are AlphaGenome predictions rather than nine billion laboratory experiments.
The resulting dataset, the presenter reports, occupies one petabyte, contains predictions for more than nine billion genetic variants, and is over 30 times larger than the AlphaFold database. He explains its use through a specific workflow: a researcher looks up a mutation in the existing Atlas and retrieves predictions about effects such as gene regulation or protein production. The model need not be run from scratch for every lookup. He also introduces a measure called AVI, describing it as an impact score for quickly identifying variants worth further investigation.
The presenter gives two examples of reported use. Researchers found 22% more genetic associations in UK Biobank data, he says. The Atlas also helped support the solution of a previously unsolved rare disease case. He sees the central change as turning a large body of previously difficult-to-interpret genetic information into a searchable database. The Atlas is accessible through Explore AlphaGenome Atlas on the project page, followed by searches across more than nine billion mutations.
LingBot-World 2.0: a continuously generated controllable world
LingBot-World 2.0 is presented as a real-time interactive world model. It continuously creates a virtual environment that a user navigates with keyboard presses. In the examples, the presenter sees higher quality than in earlier open world models, describing the surroundings as more detailed, coherent, and higher resolution. Text prompts can also add events or effects. Another addition is an agent mechanism for NPCs: other characters can move through the environment and behave or respond in different ways.
The project team, as quoted by the presenter, claims continuous interactive world generation for more than an hour. The real-time system is said to reach 720p at up to 60 frames per second. It uses an Alibaba video generator. The presenter explains that the world is generated in chunks, allowing real-time streaming. Previous information is cached, which he describes as a form of scene memory.
Code and instructions for downloading and local execution are available. The presenter distinguishes two variants with different priorities: a larger 14-billion-parameter version produces higher quality, while a 1.3-billion-parameter version is considerably faster. Their distinction is described as the larger model’s quality versus the smaller model’s speed.
Isaac 0.5: perception, prediction, and action in one model
The presenter calls Isaac 0.5 an open robot foundation model designed for a more general understanding of its surroundings and next actions. Its inputs include images, video, language instructions, the robot’s current state, and even previous actions. From that information, it can predict the next action or answer questions, according to the account. The separately listed capabilities include locating objects, predicting what the world may look like next, and directly generating robot movements. Image understanding is therefore presented as part of a broader set of tasks.
The account gives 36 billion parameters. Training data comes from more than 35 robotic systems, which the presenter connects with applying or transferring the model to different robot types. He reports 100,000 hours of robot experience and roughly one million hours of general video. The distinguishing approach is joint training of video understanding, spatial reasoning, future-state prediction, and physical actions within one backbone. He explains the idea of joint training as allowing knowledge acquired from video and spatial tasks to help the robot choose actions and understand the physical world.
The model has been open sourced, and the project page leads to download and execution instructions.
WorldSculpt: reconstruction with individually editable objects
WorldSculpt takes multiple scene images or a video and turns them into a complete 3D scene containing separate objects. The presenter emphasizes the difference from reconstructions that resemble the original surroundings but fuse everything together. WorldSculpt reconstructs each object individually while positioning it within a shared 3D world. Each can then be moved, resized, or edited as a separate part of the scene. The described result concerns the structure of the reconstruction as well as its visual resemblance.
An additional example converts a Gaussian splat world into a scene with actual 3D meshes. He reports that the project has been released, with code and download and execution instructions in the repository linked by the Code button.
FIRE3D: meshes and materials for a simulation-ready scene
FIRE3D also reconstructs a three-dimensional scene with separate objects from ordinary photos or video. Its input may be one image, multiple images, or a video recording. The presenter describes the result as simulation ready: each object becomes a complete separate mesh with material information. It can be moved, edited further, or added to a robotic simulation.
The presenter gives specific speed figures: a simulation-ready scene can be recreated in under one minute, and up to 16 objects are processed in parallel on one GPU. These are reported project figures. Code has been released with local execution instructions, and training code is also available. The presenter explicitly highlights that addition.
AuK: generating speech, precise edits and changes to its sound
Tencent’s audio model is called AuK and compared by the presenter with Nano Banana for speech. The first demonstration generates speech from a description of the desired voice and supplied dialogue. The example speaks about someone who always promised to return and kept his word until now. The presenter describes the voice as sad and grief-filled. Next is zero-shot voice cloning: a few seconds of a reference voice can be uploaded and used to speak different text. The reference and generated output are played in sequence.
The next demonstrations edit existing audio by adding or removing words. In the first short clip, a person says mamba out. The presenter inserts never and plays the edited version, describing the insertion as seamless. A different example removes an entire sentence and plays what remains. The presenter again emphasizes a natural transition after the edit.
A longer example concerns age and change. The original contains a remark about answering a question in one’s twenties, a statement that change will be constant, and a response praising the question. The edit inserts the idea that the most important lesson is to accept changes between the statement about constant change and the next response. The presenter plays the versions before and after the insertion and finds the transition very natural. This demonstrates adding a complete phrase within a conversation rather than a single word in a short clip.
The next group concerns recording quality and emotion. A noisy, messy clip is supplied for noise removal or voice enhancement, followed by a separate audio resolution and quality improvement example. Input and output are played consecutively in both cases. Another existing sentence about daily invitations to social occasions and how pleasant they were is first given a sad tone and then an angry one. The presenter thereby demonstrates changing the emotion of recorded speech while retaining the sentence’s basic content.
Timbre can also change: an original male voice becomes a female voice described as clear and bright. The example plays a news sentence about harassment of children at Flinders Street Station and a $700 fine; it is demonstration audio rather than a separate story in the overview. Another clip about the eastern coast becomes a whisper. Additional listed operations include adding or removing laughter and breaths. The GitHub repository supplies local execution instructions, and the presenter gives a model size of about 6.12 GB. From that he expects it to fit most consumer GPUs and calls it one of the most flexible speech editors he has seen.
DeepSeek-V4.1-Flash: a new architecture and mixed leaderboard positions
In the September 13 review, the presenter considers DeepSeek-V4.1-Flash a much larger change than a 0.1 version increment suggests, saying its design differs substantially from the previous V4. He begins with a score of 74.2 on DeepSWE 1.1, which he interprets as placing Flash above GPT-6 Astra at the top of the table. He immediately qualifies this: the DeepSWE team has not officially added the model, so its eventual placement remains to be seen. He then describes CyberGym and AutomationBench performance as state of the art, specifically saying Flash surpasses GPT-6 Astra Max on AutomationBench.
The supplied specifications describe a 552-billion-parameter mixture of experts with only about 8–16 billion parameters active during use. The presenter explains it as a team of experts of which only some participate in a particular task. He says the architecture is new, with separate encoder and decoder components, contrasting it with most modern language models, which he describes as decoder-only. He notes other architecture changes but focuses on the separate input-processing and answer-generation components.
The later comparisons give a less uniform picture than the initial leadership claim. On LiveBench by Abacus AI, the presenter places the new DeepSeek first among open models, a few points below GPT-6 Astra and Claude Fable. In VALS, an index covering various knowledge-work tasks, he again calls it the leading open model, slightly above Kimi K3 and GLM 5.3. He specifically says the difference is not significant and these models are effectively tied. In Artificial Analysis’s current intelligence index, however, DeepSeek remains behind Kimi K3 and GLM 5.3.
API use is reported at 217 output tokens per second, which the presenter compares with over three times GLM’s speed and roughly seven times Kimi K3’s. He describes the cost as much lower than other leading open models and the closed GPT-6 and Fable models, without supplying exact prices. The model is open, but the original download occupies about 510 GB. Community-customized and quantized versions already exist; the smallest displayed GGUF Q1 takes 106 GB. He points to the API for those without local hardware and considers DeepSeek the most cost-efficient frontier model.
OpenAI and Navier–Stokes: a result with external forcing
OpenAI announced a result concerning Navier–Stokes. The presenter explains the underlying question using ordinary movements of water and air. The equations describe liquids and gases, from water swirling in a cup to air moving through a room. The question is whether conditions exist under which that description breaks down. He associates it with the Millennium Prize problems and says it has remained unsolved for roughly 90 years. In his account, OpenAI used an internal model the company claims is substantially more capable than GPT-6 Astra.
The presenter explains the proposed result through a whirlpool under very specific conditions. The vortex keeps stretching, with speed growing without limit in a finite time. He calls this a singularity or blow-up and interprets it as evidence of a possible breakdown under certain conditions. He says the work used 10,000 concurrent agents for 88 hours. In a displayed comparison, the internal model performs better on open mathematical problems. He contrasts those 88 hours with the decades he attributes to human efforts and regards the announcement as a major mathematical breakthrough.
The presenter then introduces substantial qualifications and distinguishes versions of the problem. He associates the simpler level with Euler equations, which omit viscosity. He explains viscosity as a liquid’s thickness, using honey as more viscous than water. Another distinction is external forcing. Stirring water is his example: it guides the situation and helps establish the required conditions. Around the same time, he says, mathematicians, one associated with Anthropic, solved the easier forced Euler version.
OpenAI’s announced result includes an external force while retaining the fluid’s viscosity. In his explanation, that makes it harder than the version without viscosity. He explicitly says the unforced problem remains unsolved. The remaining question is whether water or gas can naturally reach a singularity without someone stirring it or directing it into the required vortex. The unforced Navier–Stokes problem remains open. He nevertheless continues to consider it an important breakthrough and expects further acceleration in such achievements.
Show-Harness: a vision model chooses robot actions
Show-Harness asks whether an ordinary image-understanding model can control a robot. An existing vision-language model receives a list of meaningful movements: move left, rotate, move forward, or manipulate something. It examines the scene, reasons about the next step, and selects an action from that set. The presenter emphasizes that it remains a vision-language model rather than a direct robot action model. Its commands therefore pass through a robot-specific action interpreter that converts them into actual movements.
According to the presenter, the scheme works and success can reach 100% when the model receives meaningful action names. The approach is then packaged as a harness accepting different vision-language models, including Gemini, GPT, and open options called Qwen or GLM. A GitHub repository with setup instructions is available. He also states a practical condition: a real robot is needed; the model and code alone do not replace the hardware.
Edge0: streaming only the experts that are needed
Edge0 is presented as a way to run large language models on devices without enough memory for their entire weights. It targets mixtures of experts and is explained using a model called Qwen 3.5 35B. The presenter compares the structure with a team of specialists: a request routes to the relevant experts, such as a math specialist for a math question. Only three billion of its 35 billion parameters are active during use, he says. Ordinary execution nevertheless requires enough memory for the whole model. Edge0 instead streams in only the experts needed for the current question, making memory use depend more on the active portion.
The model is also compressed to Int4, a more compact way to store its numbers. The presenter acknowledges possible quality loss and introduces Recover-LoRA as a way to recover much of it. The displayed table reports peak active memory of 2.9 GiB (about 3.1 GB) for the 35-billion-parameter version and 1 GiB (about 1.1 GB) for the smaller eight-billion-parameter version. Both were tested on a Mac mini M4 Pro with 24 GB of memory and short contexts; the figures describe active model usage rather than all the computer’s memory. The smaller version is based on Inclusion AI’s Ling 3.0 Tiny. The approach combines selective streaming, compression and quality recovery. The page supplies download and Mac execution instructions.
Real-SWE: company requirements extend beyond writing code
The presenter contrasts Real-SWE with software benchmarks called SWE-bench and DeepSWE, which he considers rapidly saturated as models solve a large share of their tasks. The new test examines work on actual company software: agents receive private code and must make a change relevant to the business. Examples include fixing invoice tax calculations and migrating customer accounts. He locates the difficulty in understanding the particular product and rules scattered across different systems. A common failure is missing a requirement. Writing code is therefore only part of the job: the agent must determine everything the change requires and integrate it correctly into the existing product.
Results remain below 40% even for Fable and GPT-6. In the leaderboard cited on September 13, Fable 5.1 is first at 38.8%, GPT-6 Astra second at 33.8%, Gemini 3.8 Flash third, and the best open model, GLM 5.3, fourth. The presenter does not state percentages for the last two. He also cautions that the four are not significantly different and are effectively statistically tied for first.
Suno V6: song generation and edits to individual parts
According to the presenter, Suno V6 increases control both over creating a song from scratch and over subsequent precise edits. New music can start from a text prompt or an audio reference. The focus is specific instruction-based changes: replacing a single lyric without affecting the rest, combining vocals from one song with an instrument from another, or isolating a guitar riff and building a beat around it. A music demonstration follows.
The presenter distinguishes three released variants. Standard V6 is available on paid plans and intended for reliable, precise results. V6 wild is also paid but described as less predictable and more varied, for unexpected ideas. V6 mini is available to everyone, including the free plan; it is much faster but lower quality and less precise than V6. He then notes Suno’s closed and paid nature and says restrictions have been added, including download limits.
YuE2: a musical plan before the finished song
The next open music generator is YuE2. The presenter explains its distinctive intermediate musical plan. Rather than moving directly from a text prompt to final audio, the model first creates something closer to a score containing melody, rhythm, chords, and structure. The plan can be edited before the model renders a complete song with vocals and accompaniment. He describes the potential to change melody, lyrics, tempo, and arrangement at a structured musical level rather than only revising prompts and hoping for the desired outcome.
Several music samples introduce the output. One English vocal addresses John and refers to rhythm, boogie-woogie, and wanting to dance. Another includes All Saints’ night, visits to neighbors, remembrance of the dead, sweets, and joy mixed with grief. These samples demonstrate complete vocal songs. A separate capability is covers: the specific example supplies Jingle Bells for a new rendition in a minor key.
From the displayed benchmarks, the developers, as quoted by the presenter, consider YuE2 the best open model, ahead of Minimax Music 3 and ACE-Step 1.5, and even claim better song quality than Suno V6. Project materials and a GitHub repository with download and local execution instructions are available. The presenter gives a model size of 7.3 GB and consequently expects it to run on most consumer GPUs.
ChatGPT for Financial Services: data, models, and traceable sources
ChatGPT for Financial Services is described as GPT-6 Astra combined with financial datasets and tools for banking research. The presenter lists company investigation, financial model building, and turning the results into spreadsheets, documents, or presentations. Named data providers include PitchBook and CrunchBase. A separate feature is tracing figures to a specific table or passage so an analyst can check their origin. Firms can also supply their own templates for output in their usual format.
The presenter suggests the features could automate considerable financial work; this is his assessment of their possible use. He immediately qualifies availability: the product is not accessible to everyone. Access must be requested by contacting sales through the page.
UMR: transferring human motion while respecting the robot’s body
UMR stands for Unified Motion Retargeting and translates human movement into movement a humanoid robot can learn from. The presenter explains why direct copying is problematic: humans and robots differ in proportions, joints, and movement limits. UMR represents both body surfaces as collections of 3D points and learns the correspondences between them. He explains it as matching the shape of a human pose to the robot’s body while respecting that body’s constraints.
According to the presenter, the approach helps preserve contact with objects and the environment. Simple examples include picking up a ball or sitting on a chair. More complex movements include spin kicks, climbing stairs, and playing tennis. He describes the outcome as reusing human motion data across robots without manually redesigning each movement. The code is available with local setup instructions in the repository.
Unitree UnifoLM-WLA: hand actions and whole-body coordination
The open Unitree model is called UnifoLM-WLA and has six billion parameters. It takes the robot’s visual input, an instruction, and current-state information, then produces actions. According to the presenter, one model covers 64 tasks: 54 tabletop tasks and ten requiring whole-body coordination. It supports two-finger grippers and different five-finger robotic hands, so the task set is not limited to a single hand type. He emphasizes the combination of a broad task list and a model he considers small.
The key training idea is predicting which parts of the scene will change during an interaction. Folding a towel is the presenter’s example: the robot learns to anticipate the relevant motion within the scene. That prediction connects to an action-generating component producing the next robot movements.
Demonstrations include loading laundry into a washing machine. In another sequence, the robot picks up a bottle and discards it, removes the trash bag, walks to a garbage can and disposes of the waste. The presenter explicitly calls this whole-body coordination. Another example combines object manipulation with walking around a kitchen. Models and the training dataset have been released. Through Models, the model is available for download at under nine gigabytes. He suggests this could allow local offline execution within a robot.
MiniCPM5-2B: a small model and its training data
Another new model is OpenBMB’s MiniCPM5-2B. The presenter introduces it as a small model for offline use on edge devices. It has two billion parameters. The displayed benchmarks cover code reasoning, mathematical reasoning, instruction following, and general knowledge. He considers it state of the art among similarly sized models and says its average performance exceeds even Qwen 3.5 4B and other small models.
The model is released with the high-quality training dataset behind it. The presenter says that dataset is also accessible for training a small model of one’s own. On the Hugging Face page, he reports a model size of about five gigabytes and expects that size to fit most consumer devices.



