Being-H0.8: A Latent Tactile World-Action Model at Scale

Date:2026.07.28

Views: 59

On July 28, BeingBeyond officially released Being-H0.8, the world's first latent tactile world-action model.


Being-H0.8 is the first to introduce touch into large-scale foundation model pretraining, unifying vision, touch, actions and future state changes in a shared latent space. This enables robots to understand not only what they see, do and touch, but also how the world changes as a result, supporting a more complete understanding and prediction of physical interactions.



BeingBeyond was among the first teams in the industry to propose training general-purpose embodied foundation models on large-scale egocentric human video. Through successive iterations from Being-H0 to Being-H0.8, the company has assembled more than 500,000 hours of raw egocentric video and established partnerships with dozens of data providers.


Data scale, however, is only the starting point. What truly determines model capability is whether the data accurately and effectively represents actions, contact and interaction.


To achieve this, BeingBeyond has built a full infrastructure stack spanning data pipelines, model pretraining, post-training, evaluation and on-device deployment. This system supports the continuous accumulation of data at scale and ongoing improvements in model capabilities, helping robots operate in increasingly open and complex real-world environments.


Building on this infrastructure, BeingBeyond has developed UniHand 3.0, a high-quality human-interaction dataset. It covers natural interactions with objects, long-horizon tasks, a rich range of hand motions and diverse environments, providing embodied foundation models with human interaction experience at scale and at high quality.


unihand.webp


UniHand 3.0 is currently China's largest human-video dataset. It also marks the first time a team anywhere in the world has systematically shared data provenance and other key information on the use of human video at this scale, offering a practical reference for large-scale use of human video in embodied AI.


"Being-H0.8 marks an important milestone," said BeingBeyond founder Zongqing Lu. "It advances model capabilities from vision to touch, and from observing the world to understanding interaction. Starting with Being-H0.8, a new generation of embodied foundation models—capable of visual understanding, tactile interaction, action generation, world prediction, and cross-embodiment generalization—is emerging. With an unprecedented learning paradigm, BeingBeyond is shaping the future of general-purpose embodied foundation models. "


overview.webp


https://research.beingbeyond.com/being-h08


The First Embodied Foundation Model to Incorporate Touch During Pretraining


Being-H0.8 extends the latent World–Action Model introduced in Being-H0.7 from visual prediction to tactile-aware interaction learning.It is also the first embodied foundation model pretrained on large-scale egocentric human video together with tactile information.


Adding touch allows the model to perceive contact states, changes in force and object constraints that are difficult to observe directly through vision. This enables a range of contact-rich tasks that are difficult to perform reliably using vision alone.


Slow-Fast Action Expert.webp


To make full use of high-frequency tactile feedback, Being-H0.8 introduces the Slow-Fast Action Expert. It adds a lightweight fast-update stream to the standard mechanism for generating a complete action chunk. The system first computes and caches the world-action context at the beginning of each action horizon. At a small number of anchors within that horizon, it then incorporates the latest observed proprioceptive state and tactile feedback to dynamically generate or revise the action segment being executed.


This mechanism preserves the overall planning capability and computational efficiency of action-chunk prediction while enabling rapid adjustments based on real-time contact feedback. It provides greater stability and responsiveness for contact-rich manipulation such as precision grasping, insertion and removal, twisting and assembly.


TactoHand Makes Tactile Data Scalable


Touch is one of the most critical sensory modalities for dexterous manipulation, yet it has long been missing from much of the available data. In many tasks, people rely on touch to determine whether an object is in contact, whether it is slipping and whether the applied force is appropriate. Touch directly conveys contact location and extent, force conditions and the spatial relationship between the hand and the object. Vision alone often cannot provide this information reliably.


Existing tactile sensors and data-collection systems, however, are generally expensive. Collection also depends heavily on specialized hardware, tactile gloves and controlled experimental environments. Collectors are often constrained to laboratory settings, making it even harder to capture a sufficiently diverse range of real-world objects, tasks and interactions. As a result, tactile data has remained difficult to scale to the volumes required for foundation-model pretraining.


1789033052107212.gif


To address this bottleneck, BeingBeyond developed TactoHand, a dense tactile pseudo-labeling system for large-scale unlabeled human video. Without requiring collectors to wear tactile gloves or use additional sensors, it infers contact probabilities and continuous proximity at densely sampled spatial points during hand–object interaction in ordinary human video. This adds tactile supervision to vast video collections at near-zero marginal cost.


Using TactoHand, BeingBeyond has, for the first time, extended tactile information to more than 500,000 hours of egocentric human video and used it for embodied foundation-model pretraining. For the first time, this turns touch from an expensive, scarce laboratory signal into a pretraining resource that can be produced and used at scale.


TopoHand Connects Human Hands, Dexterous Robot Hands and Parallel Grippers


Robots requires more than increasing data volume. It also requires resolving action-space inconsistencies across heterogeneous embodiments, including human hands, dexterous robot hands and parallel grippers.


In Being-H0.5, the team first introduced the Unified Action Space (UAS). It maps control variables such as end-effector pose, gripper opening and closing, finger joints and robot state into a shared representation through predefined semantic slots, supporting more than 30 heterogeneous robot embodiments. The first-generation UAS, however, relied primarily on action slots and their masks for cross-embodiment compatibility. It did not yet unify human hands and different robot end effectors at the level of kinematic structure.


Being-H0.8 introduces TopoHand, the second-generation unified action space, providing a common kinematic interface for human hands, dexterous robot hands and parallel grippers. TopoHand uses a fixed-topology "screw-hand" representation with 20 semantic keypoints and 20 canonical joint variables. Its shared coordinate frame is centered on the wrist: the x-axis points toward the projection of the middle fingertip onto the palm plane, the z-axis follows the palm normal, and the y-axis completes the right-handed frame.


topohand.webp


This representation keeps embodiment-specific joint names, mechanical structures and mesh topologies outside the policy interface. Instead of directly relying on the native joint definitions of a particular robot or human-hand model, the model learns manipulation patterns within a shared semantic topology and kinematic space.


Compared with the first-generation UAS, TopoHand goes beyond compatibility among control variables to bridge structural differences between human hands, dexterous robot hands and parallel grippers. This enables more efficient transfer of manipulation priors from large-scale human video to different robot embodiments, significantly improving the efficiency and scalability of cross-embodiment pretraining.


In roughly one year, BeingBeyond has reached several key milestones in large-scale egocentric human video pretraining, successively releasing the world's first embodied foundation models pretrained at each of four scales: thousands, more than 10,000, 200,000 and 500,000 hours of human video.


4e7f4fa260d435ae6c5e887f20b4a712.webp


These successive model generations address the central data challenges facing general-purpose embodied foundation models at different stages: How can egocentric human video be used? How can it be used for model pretraining at scale? And how can high-quality information be extracted from data to continually improve models?


The next question is more consequential: What training paradigm can accelerate the development of general-purpose embodied foundation models?


As it moves toward a new 1.0 era, BeingBeyond is committed to developing its own answer.


About BeingBeyond


BeingBeyond focuses on developing and applying general-purpose embodied foundation models. It was among the first companies in China to propose a general-purpose model framework trained on large-scale human video data, and the first AI startup in China to introduce a natively latent-space world model. With a mission to bring humanoid robots out of the laboratory and into everyday life, BeingBeyond is committed to solving the core technical challenges of embodied AI and leading the transformation of humanoid robotics.