Robots have traditionally needed extensive programming, task-specific training and carefully controlled workspaces. A newer approach called physical AI aims to make machines more adaptable by helping them understand the physical world, learn from demonstrations and adjust their actions when conditions change.
One recent example is Skild AI’s S1 robot foundation model. The company says an operator can show S1 a single video of an unfamiliar task and the robot can attempt the task without changing its model weights or completing separate task-specific post-training. Reported demonstrations include plant potting, pour-over coffee, pancake cooking and multi-step industrial assembly.
This does not mean a robot can safely master any job after watching one clip. It shows how video prompting and in-context learning may reduce the time needed to teach certain machines new behaviors.
Last reviewed: October 2, 2026
What Is Physical AI?
Physical AI combines artificial intelligence with machines that sense and act in the real world. Examples include industrial robots, warehouse systems, autonomous vehicles, drones and humanoid robots. Unlike a text-only assistant, a physical AI system must connect perception, reasoning and movement while obeying real constraints such as gravity, friction, object weight and human safety.
A typical system uses cameras or other sensors to understand its surroundings, a model to interpret the goal, and control software to convert decisions into movement. It may also use simulation and synthetic data to practice situations that would be slow, expensive or unsafe to reproduce repeatedly with hardware.
Physical AI is related to multimodal AI because robots often combine video, images, language, position data and touch. The difference is that mistakes can have physical consequences rather than remaining on a screen.
How Can a Robot Learn From One Video?
Skild describes S1 as an in-context learner for robotics. In-context learning means the demonstration becomes part of the model’s current input, much like an example included in a prompt for a language model. The example guides behavior without permanently updating the model’s weights.
According to Skild AI’s technical introduction to S1, a demonstration can come from a different viewpoint, scene or robotic body. The model must infer the demonstrator’s intent, identify useful object relationships and track progress through a sequence of actions.
1. A Person Demonstrates the Task
An operator records a video showing the desired result and the order of actions. For a delicate or long task, a visual example can communicate details that would be difficult to describe completely with text.
2. The Model Interprets Intent
The robot must distinguish essential actions from irrelevant details in the recording. It identifies objects, estimates spatial relationships and infers what each step is trying to accomplish.
3. The Demonstration Becomes Context
The video guides the pretrained policy at execution time. The same model weights are used rather than creating a new specialized model for that task.
4. The Robot Maps the Example to Its Body
A human hand and a robotic gripper move differently. The system must translate the demonstrated goal into actions that suit its own joints, reach, sensors and tools.
5. It Executes and Adjusts
During execution, sensor feedback helps the robot follow the sequence and respond to small changes. A capable system may notice a misplaced part, retry a grasp or continue after a minor disturbance.
6. Humans Validate the Result
Operators check accuracy, safety and consistency before trusting the behavior in production. A successful demonstration is only the beginning of deployment testing.

What Skild AI Reports About S1
Skild says S1 can complete unseen tasks lasting up to ten minutes using one video prompt and no task-specific fine-tuning. Its examples combine multiple manipulation skills rather than repeating one simple movement.
An NVIDIA overview of the collaboration describes an industrial workflow in which a robot installs components, fastens 16 screws and adapts to disturbances across multiple steps. NVIDIA says Skild uses its infrastructure for synthetic data, simulation, training and deployment.
These results are promising, but company demonstrations are not the same as independent evaluation across different robots, workplaces and failure conditions. Readers should treat them as evidence of a developing capability rather than proof of a general-purpose household or factory robot.
Video Learning vs Traditional Robot Programming
| Approach | How a new task is added | Main advantage | Main limitation |
|---|---|---|---|
| Rule-based programming | Engineers write explicit motion and control logic | Predictable in stable environments | Changes can require substantial reprogramming |
| Task-specific machine learning | Collect demonstrations and fine-tune a model | Handles variation better than fixed rules | Data collection and training can be expensive |
| Language prompting | Describe the goal in text | Fast and convenient for simple instructions | Hard to express delicate physical technique |
| Video in-context learning | Show a demonstration as part of the prompt | Can communicate sequence, motion and preference | Still depends on prior training, perception and safe execution |
These approaches can be combined. A factory may keep verified control rules for hazardous motions while using video demonstrations to configure flexible parts of a workflow.
Where Physical AI Could Be Useful
- Manufacturing: adapting assembly work when product designs or component positions change.
- Warehouses: handling a wider range of packages, shelves and packing procedures.
- Inspection: navigating facilities and checking equipment in varied environments.
- Agriculture: identifying crops and performing tasks that change with weather and growth.
- Healthcare support: moving supplies or assisting with carefully limited, supervised routines.
- Food preparation: learning multi-step processes where objects and timing vary.
- Home assistance: completing selected chores after extensive safety validation.
The strongest early uses are likely to be structured workplaces where tasks repeat but not identically. Organizations can control the environment, measure errors and place humans nearby to supervise.
Why Physical AI Is Difficult
Language models can generate another answer after an error. A robot may drop a component, damage equipment or injure someone. Physical deployment therefore demands higher reliability and better failure detection.
- Limited robot data: real-world demonstrations are costly and slower to collect than text or images.
- Reality gaps: behaviors that work in simulation may fail with real lighting, friction or object variation.
- Long sequences: a small mistake early in a ten-minute task can disrupt every later step.
- Hardware differences: a policy must account for different cameras, grippers, reach and strength.
- Unusual situations: workplaces contain people, clutter, damaged parts and events absent from training.
- Maintenance: sensors drift and mechanical parts wear, changing the conditions the model expects.
The Stanford AI Index 2026 documents progress in robotics and autonomous motion while also showing why evaluation across different platforms and environments remains important.
Safety Requirements for Robots That Learn
Before deploying a video-taught robot:
- Restrict the robot’s workspace, speed, force and tool access.
- Test the demonstration for ambiguity and missing safety steps.
- Use collision detection, emergency stops and physical guarding.
- Evaluate normal operation, disturbances and deliberate edge cases.
- Require human approval before high-risk actions.
- Log sensor data, model decisions, interventions and failures.
- Revalidate after software, hardware or workplace changes.
- Train operators to recognize unsafe or uncertain behavior.
- Keep a reliable manual fallback and incident response plan.
The same principles used for AI agent safety—limited permissions, monitoring and human approval—become even more important when software controls machinery.
Does One Video Really Mean One-Shot Learning?
The phrase can be misunderstood. The robot is not learning from nothing. S1 is a pretrained foundation model built from a broad collection of robotic experience. The single new video supplies task context at execution time.
A better analogy is a skilled worker watching one demonstration of a new assembly. The worker can understand it quickly because years of related knowledge already exist. Likewise, a physical AI model depends on prior training that taught it about motion, objects and general robot behavior.
The critical question is how far the new task falls outside that prior experience. A video cannot guarantee success when the required tool, motion or environment is completely unfamiliar.
Physical AI and AI Agents
An AI agent plans and uses tools to reach a goal in software. Physical AI extends a similar idea into machines that can perceive and act. A warehouse robot might receive an order, plan a route, identify a package, grasp it and report completion.
The combination creates more capability and more risk. An incorrect software action may alter a file; an incorrect physical action may create immediate damage. Organizations need clear boundaries between planning, approved motions and safety-critical control.
What Physical AI Means for Workers
More adaptable robots may reduce repetitive programming and make automation practical for shorter production runs. Workers may spend more time demonstrating tasks, checking quality, managing exceptions and maintaining systems.
The impact will vary by workplace. Some roles may change or decline, while new work appears in robot operations, safety, data collection and integration. Responsible deployment should include worker consultation, training and transparent measurement of both productivity and safety.
Frequently Asked Questions
What is physical AI?
Physical AI uses artificial intelligence in machines that sense, reason and act in the real world, including robots, autonomous vehicles and drones.
Can a robot learn a task from one video?
Some pretrained robot models can use one video as in-context guidance for selected new tasks. Success depends on prior training, task difficulty, hardware and environmental conditions.
Does video learning update the robot’s model?
In Skild AI’s reported S1 workflow, the demonstration guides execution without changing model weights or completing task-specific post-training.
Is physical AI the same as a humanoid robot?
No. Humanoid robots are one form of hardware. Physical AI can control many machine types, including robotic arms, vehicles and drones.
What is the biggest physical AI risk?
The main difference from screen-based AI is that errors can cause physical harm. Safe systems need strict limits, independent testing, emergency controls and accountable human oversight.
From Showing to Doing
Video in-context learning points toward robots that can be instructed by demonstration instead of lengthy reprogramming. Skild AI’s S1 examples show why the idea is attracting attention, but broad reliability will require more independent testing across machines, tasks and real workplaces.
Explore practical AI tools on Unlimited AI, and remember that physical systems require stricter testing than software-only assistants because their actions affect people and equipment directly.



