Skip to content

Verify evals on Papers with Code #118

Description

@NielsRogge

Hi,

Niels here from the open-source team at Hugging Face. Congratulations on your work!

I've made the paper and 10 verified paper-native evaluations available on Papers with Code.

The paper has results on Agents, Coding Agents, and World Knowledge task pages.

The GLM-5 result currently ranks second on τ²-Bench.

Would it be possible to verify these results and let me know if any score, model name, benchmark protocol, or openness metadata should be corrected? The imported rows are tied to the paper or its official release artifacts; comparison-table baselines were not added.

You can also edit the task, methods, project page, and GitHub URL directly from the paper page using your Hugging Face account.

If you'd like to showcase the results in your repository README, you can copy these live leaderboard badges (or use the “Copy PwC badge” button in the Results section):

Papers with Code: #2 on τ²-Bench

Kind regards,

Niels

Metadata

Metadata

Labels

No labels
No labels

Type

No type

Projects

No projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions