Article 12 Evaluating LLMs Under Production Parity: A Replay Pipeline for Safe Model Swapping in Conversational Agents