About the role
You will lead the development of the AI systems powering a humanoid robotics platform, with a focus on Vision-Language-Action (VLA) models and world models. You will be responsible for building the core intelligence layer that enables humanoid robots to perceive their environment, reason about complex tasks, and generate reliable, autonomous actions in real-world settings.
Combining deep research expertise with practical engineering judgment, you will drive research and development efforts across embodied AI, multimodal learning, simulation-driven training, and scalable data generation. Working closely with robotics engineers, you will translate state-of-the-art AI breakthroughs into production-grade systems integrated into the full humanoid robot stack for real-world deployment in factories and warehouses.
This position is in an early-stage startup environment. You will have significant ownership and autonomy, so you should be a self-starter who can operate independently, define technical direction, make key architectural decisions, and navigate ambiguity without relying on detailed guidance.
Responsibilities
Define and lead the AI roadmap focused on Vision-Language-Action models and world models for humanoid robotics.
Develop and improve multimodal AI models that connect perception, language understanding, reasoning, and robot actions.
Research and implement approaches to improve training efficiency, model performance, and deployment speed.
Develop strategies for data collection, synthetic data generation, simulation, and large-scale model training.
Design evaluation frameworks to measure model performance, generalization, and real-world robot capabilities.
Adapt and deploy state-of-the-art AI research into practical humanoid robot applications.
Collaborate with robotics engineers to integrate AI models with robot control and execution systems.
Build and lead the AI team as the company grows.
Main requirements
PhD in machine learning, artificial intelligence, robotics, computer vision, or a related field.
2+ years of industry experience developing Vision-Language-Action models, world models, embodied AI systems, or related multimodal AI technologies.
Strong understanding of multimodal learning, foundation models, and modern deep learning architectures.
Hands-on experience training, fine-tuning, and scaling large neural networks.
Strong Python and PyTorch skills.
Experience with large-scale GPU training infrastructure and distributed model training.
Experience developing AI systems using robotics datasets, video data, or large-scale trajectory data.
Other requirements
Experience with robotics simulation environments such as Isaac Lab, Isaac Sim, MuJoCo, or similar platforms.
Experience applying AI models to robotics, autonomous systems, or other real-world embodied applications.
Fluent English.
Nice to have
Published research contributions in AI/ML, robotics, computer vision, or embodied AI, with evidence of developing novel methods or approaches.
Experience with Vision-Language-Action models (e.g., OpenVLA, RT-2, π0, or similar approaches).
Experience developing world models, video prediction models, or generative models for embodied AI.
Experience leading AI research projects or technical teams.
