Being-H-Flash Achieves the World's First Real-Time On-Device Deployment

Date:2026.06.04

Views: 35

Over the past year, embodied AI has begun to move beyond purely reactive Vision-Language-Action (VLA) policies toward models that can explicitly reason about how the world may evolve. As benchmark performance improves, however, a more practical question is becoming increasingly important: can this future-aware reasoning run efficiently on the compute hardware actually carried by a robot?


For world-action models to move from research prototypes to large-scale deployment, they need to operate within the latency, power, memory, and cost constraints of onboard computing.


BeingBeyond is now taking a step toward that goal with Being-H-Flash, a highly efficient world-action model designed for real-time deployment on edge hardware. In BeingBeyond's internal tests, Being-H-Flash runs in real time on a roughly 100-TOPS-class accelerator while also supporting both Chinese AI accelerator platforms and NVIDIA GPUs.


The release builds on Being-H0.7, introduced earlier this year, which brought future-aware reasoning into robot policies through a latent-space world-action modeling architecture and scaled pretraining to 200,000 hours of egocentric human video. Being-H-Flash takes the next step: showing that future-aware robot intelligence can be made not only capable, but practical enough to run directly onboard a robot.


This reflects a broader shift in embodied AI. Model quality is no longer defined by benchmark performance alone. In real-world robotics, inference latency, hardware compatibility, power consumption, deployment cost, and engineering reliability increasingly matter just as much.


The First Real-Time World Action Modeling on ~100-TOPS Edge Hardware


To operate effectively in the physical world, a robot needs more than object recognition and instruction understanding. It also needs to anticipate how its actions may change the scene: where a moving object will be a moment later, how fabric will deform when grasped, or whether a container will spill when tilted.


This ability to reason about future interaction is particularly important for dynamic and contact-rich tasks.


Conventional VLA policies map observations and language instructions directly to actions. While effective, their learning is heavily dependent on robot action demonstrations, which are far more expensive and limited in scale than human video. As a result, researchers have increasingly explored world-action models that introduce predictive representations of future states into robot policies.


One approach is to perform this prediction in pixel space. Video-generative world-action models can explicitly imagine future visual observations and use them to support action generation and planning. Such approaches provide a rich representation of future scene evolution, but generating visual rollouts also introduces substantial computational overhead.


Being-H0.7 takes a different approach: reasoning about the future in latent space.

1280X1280.PNG

Rather than generating future images during deployment, the model learns a compact latent representation that captures future-relevant information for action generation. This makes it possible to retain future-aware reasoning while avoiding the computational cost of explicit visual rollout at inference time.


BeingBeyond's internal deployment tests show that Being-H-Flash reaches 30–45 inference updates per second on high-end GPUs including the NVIDIA A800 and RTX 4090, and approximately 20 updates per second on a ~100-TOPS-class edge accelerator.

The importance of this result is not simply higher throughput. Running a future-aware policy directly onboard the robot shortens the perception-to-action loop, reduces dependence on network connectivity, and provides more predictable system latency.


For dynamic tasks such as object catching, production-line sorting, deformable-object manipulation, and liquid pouring, these properties are critical. They allow world-action modeling to become part of the robot's online control loop rather than an offline planning or cloud-side reasoning component.


Changing the Compute Economics of World-Action Models


The efficiency of Being-H-Flash follows directly from the latent-space world-action modeling approach like Being-H0.7.


Many video-based world models learn scene evolution by predicting future observations in pixel space. For robot control, however, not every visual detail is equally useful.


A model generating future frames may spend substantial capacity modeling details such as fine textures, lighting variations, clothing wrinkles, or background appearance. These details can be important for realistic video generation, but they are often secondary to the information needed for control: object motion, contact, geometry, task progress, and the consequences of an action.


Pixel-space generation also adds computational cost during both training and inference.


Being-H0.7 therefore moves future-aware reasoning into a compact latent representation. Instead of reconstructing future images frame by frame at deployment time, it introduces learnable latent queries between multimodal perception and action generation.


During training, these representations are shaped by information from future observations, encouraging them to capture features that are predictive of how an interaction will unfold. At inference time, the model predicts these future-aware latent representations directly from the current observation and task context, without performing an explicit visual rollout.The result is a more compact reasoning process: the model preserves information about likely future interaction while avoiding the cost of generating the future pixel by pixel.


In other words, Being-H0.7 moves future prediction from pixel space into action-relevant latent space.


Being-H-Flash builds on this principle and pushes it toward practical deployment, reducing the compute required for future-aware robot control while maintaining the ability to reason about how actions change the world.


About BeingBeyond


BeingBeyond focuses on developing and applying general-purpose embodied foundation models. It was among the first companies in China to propose a general-purpose model framework trained on large-scale human video data, and the first AI startup in China to introduce a natively latent-space world model. With a mission to bring humanoid robots out of the laboratory and into everyday life, BeingBeyond is committed to solving the core technical challenges of embodied AI and leading the transformation of humanoid robotics.