Skip to content

[Question] High variance in evaluation results for the same model on identical route (id="2091") #233

Description

@NinoNeumann

Describe the question

Hi team,

I am encountering a significant inconsistency when evaluating my model. I ran the exact same model twice on the same case, but the results are completely different.

Test Case Details:

  • Route ID: 2091
  • Town: Town12

Evaluation Results:

  • Run 1: score_composed = 21.14 (Failed)
  • Run 2: score_composed = 100.0 (Completed)

Questions

  1. Is this level of variance normal for this evaluation framework?
  2. Is the randomness in the environment (e.g., NPC behaviors, traffic manager, spawning) expected to cause such a huge gap between two runs?
  3. Are there any specific configurations or random seeds I should fix to ensure reproducible evaluation results?

Thanks in advance for your help!

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions