A curated collection of publications on large language model benchmarking and evaluation, selected to support the AI Alignment workshop series. Topics span evaluation frameworks, benchmark design and critique, multilingual and domain-specific assessments, educational applications, responsible AI, and the environmental impact of LLM inference — reflecting the breadth of the field and its relevance to Dutch higher education.
How to use this reading list
Click the
DOI link in the first column to open the original
paper. Use the search box to filter by author, title, or keyword.
Last updated: 2026-07-14 — references curated for the AI Alignment workshop materials.