OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
agent benchmark natural-language-processing gui artificial-intelligence language-model cua vlm multimodal reinforcement-learnin llm large-action-model computer-use-agent
-
Updated
Jul 28, 2026 - Python