PaperScope
LIVE · 2026-10-06 05:40 UTC

The GenAI4IDN Benchmark 3.0 - a Public Tool to Assess Generative AI Tools for the Design of Interactive Digital Narratives

Hartmut Koenitz, Jonathan Barbara, Mirjam Palosaari Eladhari

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2610.05633 v1
Category
Submitted
2026-10-04

Abstract

This paper presents GENAI4IDN Benchmark 3.0, the third iteration of an evaluation framework to assess Generative AI tools for creating Interactive Digital Narratives (IDNs). Moving beyond manual testing, this iteration introduces AI-assisted evaluation through a publicly accessible web application (https://genai4idn.com), enabling the community to run benchmarks on demand, add new models, and propose new tasks. The revised evaluation framework is "blinded" to avoid model-bias, and can handle complex media such as music, videos, and full IDNs that previously required human raters. A significant addition - responding to concerns raised during ICIDS 2025 - is the addition of fact-checking and bias detection with dedicated tasks and rubrics, validated by human raters with lived experience in the depicted contexts. Findings from a diverse range of models report on maturing creative capabilities while observing runaway thinking and overzealous safety filters as limitations. Fact-checking reliably caught subtle historical inaccuracies, anachronisms, and fabricated claims while the bias rater consistently exposed structural assumptions, tropes, and marginalized group erasures.

Comment: Accepted for publication at ICIDS 2026, Bangkok, Thailand

arXiv abs page · PDF · same-day batch