Is the #1 LLM on the Leaderboard Also #1 at Toss? Building Toss’s Evaluation Framework
A large language model that proudly holds the top spot on public AI leaderboards does not necessarily deliver equally impressive performance in Toss’s actual production services. Whenever a new model is released, we often check the measured scorecards first, looking at how well it solves math problems and how smoothly it writes code. It is … Read more