
Masterclass AI Evaluation: from General Purpose to Specific Use Cases
How can you tell if an AI model is suitable for a use case? When do you know it works as intended? Leading systems are evaluated on general capabilities, not how they fit real-world problem areas, organizational values and stakeholder concerns.
Description
Private and public sector organizations alike are eager to leverage the capabilities of AI. This requires adapting general-purpose AI (GPAI) systems to work processes they are not necessarily designed or tested for. For example, a Dutch judiciary authority’s chatbot may refuse a resident’s question about who maintains the ‘privé gedeelte’ of an apartment building, because a content filter designed in English flags the words ‘private parts’ as sexual content. Alternatively, a municipal chatbot may send Mehmet - a lifelong Rotterdamer asking where to renew his passport - to the immigration service, while Daan gets the right answer immediately, since the provider didn’t test for harms specific to the Dutch context.
To deploy AI responsibly, practitioners need understanding of how existing evaluations translate to a context and where gaps arise. GPAI is evaluated on common tests known as ‘benchmarks’. But popular benchmarks are usually limited in their linguistic or cultural scope, cover limited types of user interaction, and vary wildly in their scientific quality and robustness. Appropriate evaluation often requires building tailored benchmarks from new datasets but also asks for practical validation using custom test cases to ensure systems function as intended once deployed.
Between developing risk-monitoring benchmarks for the European AI Office, a validation framework for the Dutch Judiciary’s Rechtspraak chatbot and safeguards for Dutch generative AI, Algorithm Audit has built expert knowledge in helping organizations navigate evaluation within their specific use case. In this masterclass, we distill the most valuable insights from our work in the field of GPAI evaluation. The course covers:
- GPAI under the AI Act and benchmarking for systems with systemic risk
- Prominent industry benchmarks and adaptations to GPT-NL
- Development of custom evaluation methods and validation approaches for your AI application
- Current issues in the field from a scientific perspective
- The design aspects that determine the quality of a benchmark
After attending, participants will be able to:
- Understand relevant obligations for GPAI systems under the AI Act
- Think critically about the validity and reliability of benchmarks
- Navigate popular evaluation repositories and documentation
- Assess the suitability of existing benchmarks for their own work
- Determine when custom evaluation or validation is needed and how it may look
Date
3 November 2026
Address
The Hague Conference Centre (New Babylon), Anna van Buerenplein 29, 2595 DA Den Haag
Programme
- 09:30-10:00 Doors open
- 10:00-11:15 Introduction General Purpose AI (GPAI) and testing GPAI applications
- 11:15-11:30 Break – coffee, tea and refreshments
- 11:30-12:30 State-of-the-art concepts and developments in GPAI benchmarking
- 12:30-13:30 Lunch – catered
- 13:30-14:45 Case study from Dutch public sector
- 14:45-15:15 Break – coffee, tea and refreshments
- 15:15-16:30 Hands-on exercises to build practical experience
- 16:30-17:30 Drinks – catered
Fee
- €300 in-person participation (including, lunch, drinks, refreshments)
Audience
Professionals from private and public sector who regularly work with GPAI applications, such as implementation of generative AI solutions in work processes, testing GPAI capabilities and/or working on AI policy.