
Building a Model Evaluation Practice for Production AI
Downloads an .ics file · Times are in America/New_York · Google Calendar
Most teams pick an AI model the same way they pick a coffee machine - the demo looks good or they hear about it from someone, so they assume it works. But in production, evaluation decisions directly impact cost, latency, accuracy, and user trust. Without a structured eval practice, you're flying blind when a model regresses, or a new release claims to be better. This session gives you a framework for evaluating LLMs against your workloads and operationalizing evaluation as a continuous practice, not a one-time exercise. What we'll cover: Why standard benchmarks are often misleading for enterprise use cases Building a golden dataset: how to sample, label, and version evaluation sets from real production traffic Key metrics beyond accuracy: hallucination rates, precision/recall tradeoffs, cost-per-correct-answer, and latency percentiles Running evals at scale: open-source tools (Ragas, PromptFoo, LangSmith) and using Amazon Bedrock Model Evaluation Setting a model evaluation gate: how to enforce a promotion threshold before any model goes to production Demo: running a side-by-side eval across two models on a real classification task High level schedule: 4:30 - 4:45: Arrivals 4:45 - 5:45: Presentation & Demo 5:45 - 6:00: Q&A 6:00 - 6:30: Networking Important instructions Event starts at 4:45 pm EST, but please allow at least 15 minutes for security to process registration Please ensure that your meetup profile has your full name. Both first and last name are required and we will not be able to register attendees with just abbreviations or incomplete names. There is an optional networking and Q&A event at the end of the meetup
More like this near New York
Going to Building a Model Evaluation Practice for? Ask me anything about it.
I read the organiser's pages and answer in a few seconds.
Answers are AI-generated · Privacy




