Home
Domestic Embodied Large Model Wall-OSS-0.5 Makes a Big Splash with Open Source: Zero-Shot Deployment After Pre-training
In May 2026, China’s embodied intelligence sector achieved a significant milestone with X Square Robot’s open-source release of its latest Vision-Language-Action (VLA) model, Wall-OSS-0.5. Breaking from the industry norm of pre-exam fine-tuning, this model enables zero-shot deployment on physical robots without task-specific adjustments.

Shifting from Custom Scripts to General Intelligence
The embodied intelligence field has long struggled with a hidden limitation: most models require extensive fine-tuning for specific tasks before evaluation, making it hard to distinguish between true generalization and mere scripted behavior.
X Square Robot addresses this with Wall-OSS-0.5. Pre-trained on over 20 robot types, millions of trajectory datasets, and a 90-million-entry multimodal corpus, the model was deployed directly onto real robots without task-specific fine-tuning. The team evaluated its performance across 17 complex tasks, including semantic understanding, rigid and flexible object manipulation, and precision operations.
Key Highlights: A Major Leap in Pre-training
Test results demonstrate that Wall-OSS-0.5 surpasses current expectations:
Zero-Shot Deployment: Without fine-tuning, a version with 400k pre-training steps scored above 80/100 in four of 17 zero-shot tasks. Notably, it achieved a score of 82 in the "tightening the rope" task, a flexible object challenge not encountered during pre-training.
Enhanced Fine-tuning Potential: When targeted fine-tuning is applied, Wall-OSS-0.5 shows exceptional learning efficiency. Compared to the industry benchmark π0.5, it leads by an average of 17.5 points under the same data budget, with precision tasks like precise insertion showing nearly an order-of-magnitude improvement in success rates.
Capability Evolution, Not Degradation: Experiments reveal that after intensive action training, the model’s multimodal perception capabilities remain intact and undergo a "reformative" evolution in visual positioning and reasoning.
Four Core Technologies Establish a Competitive Edge
Wall-OSS-0.5’s performance stems from four fundamental innovations:
Gradient Bridging: By injecting action supervision signals directly into the pre-training backbone, the model unifies "seeing, speaking, and acting" at the representation level.
Visual Alignment Tokenizer: This ensures each action token carries clear visual semantics, granting the model genuine "physical meaning" inference capabilities.
Action Space Supervision: Training focuses on the overall trajectory structure rather than trivial high-frequency details, significantly boosting convergence efficiency.
DMuon Distributed Optimization: Low-level system optimizations reduced heterogeneous computing costs by 100 times, making this complex training formula viable on large-scale clusters.
A New Era for Embodied Intelligence
X Square Robot has fully open-sourced the model weights, training code, and dataset interfaces for Wall-OSS-0.5.
Industry analysts note that this release redefines the development paradigm of embodied intelligence, shifting focus from "single-task success rates" to "general physical intuition transfer." For researchers and developers, this marks the start of a new phase for foundation models—defined by "reproducibility, verifiability, and challengeability"—which will significantly accelerate the deployment of general-purpose robots in complex real-world environments.
Related article
AI Agents Breach Hugging Face Sandbox With Autonomous Attack and Defense Record
Recently, Hugging Face published a comprehensive technical timeline detailing a widely reported AI agent breach. The autonomous system, developed using OpenAI models, executed roughly 17,600 operations over 4.5 days while security protocols were temp
Final 24 Hours to Exhibit at TechCrunch Disrupt 2026
Exhibit table reservations expire tonight, Friday, September 18, at 11:59 p.m. PT. After this deadline, you will no longer be able to add your startup to the Expo Hall.Tables are limited and allocated on a first-come, first-served basis. The startups
Doubao Mobile Launches Nubia NaviX Ultra, Marketed as World’s First AI Agent Flagship
Ni Fei, CEO of Navea, revealed that the next-generation Doubao smartphone, branded as "Navea NaviX Ultra," will launch in September. Positioned as the world’s first AI smart agent flagship, it marks a significant milestone. Ni noted, "From the M153 p
Related Special Topic Recommendations
Comments (0)
0/500
In May 2026, China’s embodied intelligence sector achieved a significant milestone with X Square Robot’s open-source release of its latest Vision-Language-Action (VLA) model, Wall-OSS-0.5. Breaking from the industry norm of pre-exam fine-tuning, this model enables zero-shot deployment on physical robots without task-specific adjustments.

Shifting from Custom Scripts to General Intelligence
The embodied intelligence field has long struggled with a hidden limitation: most models require extensive fine-tuning for specific tasks before evaluation, making it hard to distinguish between true generalization and mere scripted behavior.
X Square Robot addresses this with Wall-OSS-0.5. Pre-trained on over 20 robot types, millions of trajectory datasets, and a 90-million-entry multimodal corpus, the model was deployed directly onto real robots without task-specific fine-tuning. The team evaluated its performance across 17 complex tasks, including semantic understanding, rigid and flexible object manipulation, and precision operations.
Key Highlights: A Major Leap in Pre-training
Test results demonstrate that Wall-OSS-0.5 surpasses current expectations:
Zero-Shot Deployment: Without fine-tuning, a version with 400k pre-training steps scored above 80/100 in four of 17 zero-shot tasks. Notably, it achieved a score of 82 in the "tightening the rope" task, a flexible object challenge not encountered during pre-training.
Enhanced Fine-tuning Potential: When targeted fine-tuning is applied, Wall-OSS-0.5 shows exceptional learning efficiency. Compared to the industry benchmark π0.5, it leads by an average of 17.5 points under the same data budget, with precision tasks like precise insertion showing nearly an order-of-magnitude improvement in success rates.
Capability Evolution, Not Degradation: Experiments reveal that after intensive action training, the model’s multimodal perception capabilities remain intact and undergo a "reformative" evolution in visual positioning and reasoning.
Four Core Technologies Establish a Competitive Edge
Wall-OSS-0.5’s performance stems from four fundamental innovations:
Gradient Bridging: By injecting action supervision signals directly into the pre-training backbone, the model unifies "seeing, speaking, and acting" at the representation level.
Visual Alignment Tokenizer: This ensures each action token carries clear visual semantics, granting the model genuine "physical meaning" inference capabilities.
Action Space Supervision: Training focuses on the overall trajectory structure rather than trivial high-frequency details, significantly boosting convergence efficiency.
DMuon Distributed Optimization: Low-level system optimizations reduced heterogeneous computing costs by 100 times, making this complex training formula viable on large-scale clusters.
A New Era for Embodied Intelligence
X Square Robot has fully open-sourced the model weights, training code, and dataset interfaces for Wall-OSS-0.5.
Industry analysts note that this release redefines the development paradigm of embodied intelligence, shifting focus from "single-task success rates" to "general physical intuition transfer." For researchers and developers, this marks the start of a new phase for foundation models—defined by "reproducibility, verifiability, and challengeability"—which will significantly accelerate the deployment of general-purpose robots in complex real-world environments.
AI Agents Breach Hugging Face Sandbox With Autonomous Attack and Defense Record
Recently, Hugging Face published a comprehensive technical timeline detailing a widely reported AI agent breach. The autonomous system, developed using OpenAI models, executed roughly 17,600 operations over 4.5 days while security protocols were temp
Final 24 Hours to Exhibit at TechCrunch Disrupt 2026
Exhibit table reservations expire tonight, Friday, September 18, at 11:59 p.m. PT. After this deadline, you will no longer be able to add your startup to the Expo Hall.Tables are limited and allocated on a first-come, first-served basis. The startups
Doubao Mobile Launches Nubia NaviX Ultra, Marketed as World’s First AI Agent Flagship
Ni Fei, CEO of Navea, revealed that the next-generation Doubao smartphone, branded as "Navea NaviX Ultra," will launch in September. Positioned as the world’s first AI smart agent flagship, it marks a significant milestone. Ni noted, "From the M153 p











