数据集 / Qwen/AgentWorldBench

Qwen/AgentWorldBench 已完整同步

AgentWorldBench

AgentWorldBench is a comprehensive evaluation benchmark for language world models, constructed from real-world observations of frontier model trajectories on established benchmarks such as Tool Decathlon, Terminal-Bench 1.0 & 2.0, and OSWorld-Verified. Every evaluation sample is paired with a ground-truth observation obtained from real environment execution, enabling reference-grounded scoring.

AgentWorldBench evaluates world modeling quality by scoring each predicted environment observation on five dimensions — Format, Factuality, Consistency, Realism, and Quality — probing the reasoning, knowledge, and long-context capabilities required for faithful environment simulation.

For more details, please refer to the technical report and the blog post.

Benchmark Statistics

Domain Samples Avg. Turns Description
MCP 286 23.1 API server responses: tool call results, database state, service protocols
Search 458 15.5 Search engine results: URLs, snippets, rankings, page content
Terminal 354 26.7 Command-line environment: shell output, file system state, process behavior
SWE 472 28.1 IDE / code editing environment: git diff, test results, compilation errors
Android 200 37.8 Android UI hierarchy changes after touch/gesture actions
Web 200 14.2 Browser DOM state changes after user interactions
OS 200 12.7 Desktop OS state: file system, window management, application behavior
Total 2,170 22.8

Data Format

Each file is a per-domain JSONL ({domain}_test.jsonl). Each record is a single evaluation turn from a multi-turn environment trajectory.

prompt and response are parallel lists of length turn_idx, representing the full conversation history up to and including the evaluated turn. The ground-truth observation for the current turn is always the last element response[-1], while earlier elements provide context from preceding turns.

{
    "task": "terminal",
    "id": 267463494664789,
    "prompt": [
        "### Turn 1\n**Action:**\n```json\n[{\"keystrokes\": \"ls -la\\n\"}]\n```",
        "### Turn 2\n**Action:**\n```json\n[{\"keystrokes\": \"cat README.md\\n\"}]\n```",
        "### Turn 3\n**Action:**\n```json\n[{\"keystrokes\": \"mkdir output\\n\"}]\n```"
    ],
    "response": [
        "**Environment Observation:**\nroot@2b1e6f43cde5:/app# ls -la\ntotal 20\n...",
        "**Environment Observation:**\nroot@2b1e6f43cde5:/app# cat README.md\n...",
        "**Environment Observation:**\nroot@2b1e6f43cde5:/app# mkdir output\nroot@2b1e6f43cde5:/app#"
    ],
    "current_prompt": "### Turn 3\n**Action:**\n```json\n[{\"keystrokes\": \"mkdir output\\n\"}]\n```",
    "system_str": "# Role and Objective\n\nYou are a **Terminal World Model** ...",
    "turn_idx": 3,
    "total_turns": 151
}

Fields:

Field Description
task Domain identifier (mcp, search, terminal, swe, android, web, os)
id Trajectory identifier (shared by all samples from the same trajectory)
prompt List of action prompts from turn 1 through turn_idx. prompt[i] is the action at turn i+1
response List of ground-truth observations from turn 1 through turn_idx. response[-1] is the ground truth for the evaluated turn; earlier elements are context
current_prompt The action prompt for the evaluated turn (same as prompt[-1])
system_str The world model system prompt for this sample
turn_idx 1-indexed position of the evaluated turn
total_turns Total number of turns in the source trajectory

Note: Each trajectory may appear as multiple records with

9 个文件

浏览文件