
Also streaming online: Join online
Optimising Test Cases for Evaluating LLMs
Speaker: Aldeida Aleti – OPTIMA/Monash University
Aldeida Aleti is a Professor in Software Engineering at Monash University, where her research spans the intersection of Software Engineering and Artificial Intelligence. Her work focuses on advancing quality assurance methods for AI-based systems. She has published widely in top-tier venues such as ICSE, FSE, ASE, IEEE TSE, JSS, TOSEM, etc. and serves in leadership roles across major international conferences and editorial boards. Aldeida has received numerous honours for her research and service, including multiple best paper awards, best reviewer awards, the JSS Editor of the Year Award, the Dean’s Early Career Researcher of the Year Award, the Dean’s award for PhD supervision, the Humboldt Research Fellowship for Experienced Researchers, the Australian Research Council Discovery Early Career Researcher Award, etc. Aldeida serves as Senior Associate Editor for ACM TOSEM and JSS, is on the Steering Committee of IEEE ICST, and was the chair of ACM Sigsoft Junior Awards.
Abstract:
As LLMs are used in increasingly critical settings, evaluation needs to do more than score outputs. It must reveal real weaknesses with limited test budgets. This talk explores methods for optimising test cases for LLM evaluation, with a focus on improving fault detection, coverage, and efficiency. I will discuss how carefully designed test cases can better expose weaknesses in reasoning, robustness, consistency, and context handling. The talk will also highlight strategies for generating, selecting, and prioritising tests so that evaluation reflects realistic usage conditions and reveals more meaningful model behaviours.
This event is Hybrid:
JOIN VIA ZOOM – MEETING ID: 873 1557 5255; PASSWORD: 778635
Level 8, Room 8109, Melbourne Connect
SEMINAR: WED 05 AUGUST 2026 4.00PM-5.00PM (AEST, Melbourne Time)