TurnShift: Multi-Turn LLM Evaluation Robustness
Designed an NLP evaluation experiment probing whether GPT-4o judges multi-turn conversation quality holistically or locally. Constructed 8 multi-turn dialogues across 4 domains (math, coding, factual, emotional) with detailed ground-truth quality labels. Measured a 0.33-point quality delta after turn shuffling, exposing GPT-4os local-quality bias as an automated evaluation judge.