Stars
Standardize benchmark wrapping so the community can wrap various otherwise-incompatible benchmarks uniformly and use them everywhere.
Accelerating your LLM training to full speed! Made with β€οΈ by ServiceNow Research
A scalable asynchronous reinforcement learning implementation with in-flight weight updates.
A comprehensive framework to test audio comprehension of Large Audio Language Models.
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
A bibliography and survey of the papers surrounding o1
π¨βπ» An awesome and curated list of best code-LLM for research.
PyTorch native quantization for training and inference
AgentLab: An open-source framework for developing, testing, and benchmarking web agents on diverse tasks, designed for scalability and reproducibility.
π OpenHands: AI-Driven Development
Large Language Model Text Generation Inference
xLAM: A Family of Large Action Models to Empower AI Agent Systems
ππͺ BrowserGym, a Gym environment for web task automation
WorkArena: How Capable are Web Agents at Solving Common Knowledge Work Tasks?
AgentTuning: Enabling Generalized Agent Abilities for LLMs
A library for advanced large language model reasoning
[ICLR 2024] Lemur: Open Foundation Models for Language Agents
LLMs build upon Evol Insturct: WizardLM, WizardCoder, WizardMath
Code repo for "WebArena: A Realistic Web Environment for Building Autonomous Agents"
An Open-Ended Embodied Agent with Large Language Models
The benchmark to evaluate lifelong reinforcement learning algorithms
This repository contains the code of the distribution shift framework presented in A Fine-Grained Analysis on Distribution Shift (Wiles et al., 2022).
Second Order Optimization and Curvature Estimation with K-FAC in JAX.



