Highlights
- Pro
Stars
A unified VLA post-training framework for human-in-the-loop data collection, fine-tuning, and reinforcement learning.
A feed-forward 3D foundation model for reconstructing scenes from streaming data
Skill package for ML/CV/NLP paper writing, curated and adapted from Prof. Peng Sida's open notes for Codex, Claude Code, and Gemini.
NVIDIA Isaac Sim™ is an open-source application on NVIDIA Omniverse for developing, simulating, and testing AI-driven robots in realistic virtual environments.
SWE-bench: Can Language Models Resolve Real-world Github Issues?
[RSS 2026] Causal video-action world model for generalist robot control
[CVPR 2025 Best Paper Nomination] FoundationStereo: Zero-Shot Stereo Matching
Pioneering Automated GUI Interaction with Native Agents
🚀 Efficiently (pre)training foundation models with native PyTorch features, including FSDP for training and SDPA implementation of Flash attention v2.
One framework to evaluate any VLA model on any robot simulation benchmark.
A framework for efficient model inference with omni-modality models
Our inference and training framework to run on the Cosmos Models
From Vision-Language-Action Models to a Real-World Robot Learning Stack
Wan: Open and Advanced Large-Scale Video Generative Models
In-the-Wild Compliant Manipulation with UMI-FT
PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation
AR 3D object detection for iPhone with LiDAR — YOLO 2D + BoxerNet 3D lifting
Ongoing research training transformer models at scale
[ICCV 2025] Official implementations for paper: VACE: All-in-One Video Creation and Editing




