1 Introduction

If we seek we shall find

This document provides material for a community process around generating a scorecard for AI alignment.

This material may or may not be directly applicable to a specific workshop context. Its primary purpose is to provide example structures to ease the burden later of running a community process.

2 MAIN Exercises

These exercises guide you through reviewing literature, evaluating scoring mockups, and planning next steps for building an AI alignment scorecard. Work through them sequentially, iterating as needed.

2.1 Searching for relevant information

This exercise builds familiarity with the current literature landscape and the search methodology. You will prioritise papers, critique the search strategy, and iteratively refine the query to improve recall and precision.

  1. Distribute the papers from the Background section among the group based on importance or relevance to each member’s domain expertise.
  2. Select the best two papers and write a brief motivation explaining why they stand out — consider methodology, relevance to Dutch education, or actionable insights.
  3. Search online for a stronger paper on the same subject using citation trails or the databases listed in the Identification section.
  4. Read the PRISMA section above to understand the systematic review methodology before refining your search.
  5. Review and update the search query — what terms are missing? What would return more relevant results for your specific focus?
  6. Execute the refined search across the listed databases.
  7. Record the number of papers found at each stage to document recall and precision.
  8. Iterate from step 5, adjusting the Boolean string until the results converge on a manageable, relevant set.

2.2 Review the community scoring mockup

Evaluate the current scoring dashboard and scorecards to identify gaps, inconsistencies, or improvements before finalising the framework.

  1. Visit the dashboard and interact with the scoring interface. Note your first impressions: is the layout intuitive? Are the criteria clearly defined?
  2. Compare score cards across different entries. Look for patterns — are certain dimensions consistently scored higher or lower? Does the scoring align with your expert judgement?
  3. Iterate at least once from step 5 (search term refinement). New insights from the scoring may reveal missing keywords or evaluation dimensions.

2.3 Plan next step

Consolidate what you have learned and define the path forward for creating a baseline scorecard for AI alignment in Dutch education.

  1. Scope the review effort. Does the research question warrant a full systematic review, or is a rapid evidence assessment sufficient given time and resource constraints? Consider the trade-off between rigour and feasibility.
  2. Map the path to a scorecard. Outline the remaining steps: synthesise findings, define scoring dimensions, weight criteria, validate with stakeholders, and produce the baseline scorecard for community feedback.

3 Systematic Literature Review Methodology

PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) 2020 framework to ensure transparency and reproducibility in the identification, screening, eligibility, and inclusion of relevant literature.

One possible method involves a Prisma review. This approach entails conducting a systematic literature review to gather all existing evidence from scientifically analyzed research on AI alignment. By performing a systematic review, we can avoid duplicating work that has already been reviewed by a set of experts.

This method is labour-intensive, but it is a gold standard for providing an authoritative answer to a specific set of questions. This approach is reproducible and allows for a documented path detailing how recommendations were formed.

Next, we will walk through the various parts of the systematic literature review.

3.1 Identification — Database Search Strategy

By establishing clear identification criteria, we can specify the information we seek from online resources, filtering tens of millions of research papers.

Here is an example query string to start a conversation around optimization of the search.

  • Databases queried:

  • Search strategy: The search is structured into four conceptual blocks combined with AND. Each block captures a distinct facet of the research question:

    Block Key Terms Rationale
    AI / LLMs large language model, LLM, frontier model, generative AI, foundation model Captures the full range of relevant model types used in educational contexts
    Alignment & Evaluation AI alignment, value alignment, benchmark, scorecard, rubric, evaluation criteria, quality framework, readiness, certification, audit Targets existing evaluation frameworks, criteria, scorecards, and readiness tools for judging LLMs
    Education Context education, learning, pedagogy, Dutch education, Netherlands, MBO, HBO, vocational education, curriculum, competency, formative assessment Narrows the search to the educational domain with explicit Dutch focus
    Cross-cutting Themes digital sovereignty, technological sovereignty, data sovereignty, cost, affordability, sustainability, environmental impact, carbon footprint, data standard, interoperability, educational data standard Covers broader policy, feasibility, sustainability, and data-exchange dimensions relevant to adoption
  • Search terms (combined as AND between blocks, OR within each block):

    • AI / LLMs: ("large language model" OR "LLM" OR "frontier model" OR "generative AI" OR "foundation model")
    • Alignment & Evaluation: ("AI alignment" OR "value alignment" OR "capability control" OR "benchmark" OR "scorecard" OR "rubric" OR "evaluation criteria" OR "quality framework" OR "readiness" OR "certification" OR "assessment criteria" OR "audit")
    • Education Context: ("education" OR "learning" OR "pedagogy" OR "Dutch education" OR "Netherlands" OR "MBO" OR "HBO" OR "vocational education" OR "curriculum" OR "competency" OR "formative assessment")
    • Cross-cutting Themes: ("digital sovereignty" OR "technological sovereignty" OR "data sovereignty" OR "cost" OR "affordability" OR "sustainability" OR "environmental impact" OR "carbon footprint" OR "data standard" OR "interoperability" OR "educational data standard")
  • Date range: January 2023 to June 2026

  • Language: English only

  • Document types: Peer-reviewed journal articles, conference proceedings, and preprints from reputable venues

  • Supplementary search: Citation tracking (backward and forward) from key review papers and grey literature from known research groups (e.g., alignment research institutes, educational technology labs)

3.2 Screening — Title and Abstract Review

Screening filters out irrelevant papers and enables more focus on the papers that remain.

  • Inclusion criteria:
    • Studies proposing or applying an evaluation framework, benchmark, or metric for LLMs in educational settings
    • Studies addressing AI safety, alignment, bias, or governance specific to educational contexts
    • Empirical studies (quantitative, qualitative, or mixed-methods) with reproducible methodology
  • Exclusion criteria:
    • Opinion pieces, editorials, or non-empirical commentary without supporting data
    • Studies focused solely on general-purpose LLM capabilities without educational application
    • Preprints superseded by a later peer-reviewed version of the same work
  • Process: Two reviewers independently screened titles and abstracts. Disagreements resolved by consensus or third reviewer. Percentage agreement and Cohen’s kappa recorded.

3.3 Eligibility — Full-Text Assessment

Here we filter down further

  • Full-text PDFs retrieved for all records passing screening
  • Quality appraisal criteria:
    • Clarity of research question and methodology
    • Rigor of evaluation design (sample size, metrics, baselines)
    • Reproducibility: availability of datasets, code, or detailed protocols
    • Relevance to the Dutch educational context or transferability to it
  • Exclusion reasons recorded:
    • Insufficient methodological detail
    • Out of scope after full-text review
    • Duplicate or overlapping publication
    • Low quality or irreproducible results

3.4 Inclusion — Final Selection

Here we refine our focus to the most actionable content.

  • Final corpus of studies confirmed after eligibility assessment
  • Thematic saturation verified: no new framework dimensions or evaluation approaches emerged in the final portion of screened records
  • Corpus categorised by:
    • Framework type (benchmark, rubric, guideline, audit tool)
    • Educational level (higher education, professional training, etc)
    • Alignment dimension addressed (knowledge, skills, attitude, safety, bias, governance)

3.5 PRISMA Flow Diagram

The Prisma Flow diagram below illustrates the filtering process, moving from a broad collection of raw evidence gathered using wide search terms down to highly specific details. When these details are combined, they allow one to begin answering focused questions regarding the subject at hand. In this case, concerning AI alignment, this workflow facilitates a high degree of reproducibility, meaning there is substantial evidence to support any claims made.

Identification
Records identified from databases (n = ...)
Records from other sources (n = ...)
Total records identified (n = ...)
Duplicates removed (n = ...)
Screening
Records screened (n = ...)
Excluded
By title/abstract (n = ...)
Eligibility
Full-text assessed (n = ...)
Excluded
Full-text excluded (n = ...)
Included
Studies in review (n = ...)

4 Background Information

How to read this table Each reference title is listed alongside a Citation column containing the citation key. Click Details to expand description, motivation, URL, and tag for each entry. References with incomplete metadata show only the citation key until the record is finalised.

title citation_details
OpenLearnLM Benchmark: A Unified Framework for Evaluating Knowledge, Skill, and Attitude in Educational Large Language Models (Lee, Lee, et al. 2026)
Details

Description: The OpenLearnLM benchmark presents a theory-grounded framework for assessing large language models across educational dimensions of knowledge, skills, and attitude using a 124,000-item dataset. It evaluates models on pedagogical understanding and behavioral consistency, revealing distinct capability trade-offs among frontier models across these axes.

Motivation: An example of how to construct an Educational benchmark: OpenLearnLM provides an open, comprehensive framework for advancing LLM readiness in authentic educational contexts

URL: https://arxiv.org/abs/2601.13882

TAG: BENCHMARK

Are Video Models Zero-Shot Learners and Reasoners in Education? EduVideoBench, A Knowledge-Skills-Attitude Benchmark for Educational Video Generation
Details

Description: EduVideoBench, the first balanced benchmark in the education domain, grounded in the Knowledge-Skills-Attitude (KSA) framework so that pedagogical adequacy and educational safety are evaluated jointly rather than as ad-hoc quality dimensions. Across five frontier VGMs, our results show substantial room for improvement across knowledge, skills, and attitude before they are classroom-ready.

Motivation: Example of a benchmark for a specific modularity (Video) within Education

URL: https://arxiv.org/abs/2605.26918

TAG: BENCHMARK

Towards responsible AI in education: A Delphi-AHP-based framework for evaluating educational large language models (Lin et al. 2026)
Details

Description: To steer educational LLMs towards responsible and beneficial development, this study aims to construct a multidimensional evaluation framework grounded in educational theory. Initially, a preliminary pool of evaluation indicators was established on the basis of a review of the literature and pedagogical theories. The Delphi method was subsequently employed to refine the indicator structure by integrating opinions from 21 cross-disciplinary experts.

Motivation: An example of systematic process for including a community of experts in the Benchmark design process. A possible scalable approach that we can replicate.

URL: https://www.sciencedirect.com/science/article/pii/S2666920X25001742

TAG: COMMUNITY

DUMB: A Benchmark for Smart Evaluation of Dutch Models (Wei et al. 2025)
Details

Description: Dutch Model Benchmark: DUMB. The benchmark includes a diverse set of datasets for low-, medium- and high-resource tasks. The total set of nine tasks includes four tasks that were previously not available in Dutch, but limited in extent by the definition of model

Motivation: Dutch Benchmark

URL: https://arxiv.org/abs/2305.13026

TAG: BENCHMARK

Cost, Usability, Credibility, Fairness, Accountability, Transparency, and Explainability Framework for Safe and Effective Large Language Models in Medical Education: Narrative Review and Qualitative Study. (Quttainah et al. 2024)
Details

Description: The aim of this study was to identify and prioritize the enablers for developing successful LLMs for medical education. We further evaluated the relationships among these identified enablers.

Motivation: Feedback from experts and users. The study is the first to identify, prioritize, and analyze the relationships of enablers of effective LLMs for medical education The difference between expert and user opinions being that of users prefer cheaper models

URL: https://ai.jmir.org/2024/1/e51834

TAG: COMMUNITY

Benchmarking and Evaluation Framework for Large Language Models in Education (Nikiforova-Ilieva and Georgiev 2025)
Details

Description: This paper examines the functional differences and optimal implementation of leading major language models with a focus on their functional capabilities, and performance in the context of learning and analytics tasks. In addition, a relatively new but gaining local popularity language model specialized in Bulgarian is also addressed.

Motivation: Evaluating models in the cntext of a specific language. The results show that DeepSeek achieves the highest content accuracy, ChatGPT demonstrates the strongest pedagogical value, and BgGPT excels in localized educational tasks in Bulgarian.

URL: https://doi.org/10.1109/ciees66347.2025.11300167

TAG: BENCHMARK

A comprehensive review of large language models: issues and solutions in learning environments (Shahzad et al. 2025)
Details

Description: This study provides a comprehensive review of LLMs, proposing a novel theoretical framework for their ethical integration into educational settings while addressing challenges such as bias, data privacy, and adaptability through practical case studies and actionable solutions.

Motivation: Let us see if for each challenge we can define a measurable or an actionable together

URL: https://link.springer.com/article/10.1007/s43621-025-00815-8

TAG: CHALLENGES

The life cycle of large language models in education: A framework for understanding sources of bias (Lee et al. 2024)
Details

Description: … the integration of LLMs in education technology has renewed concerns over algorithmic bias, which may exacerbate educational inequalities. Building on prior work that mapped the traditional machine learning life cycle, we provide a framework of the LLM life cycle from the initial development of LLMs to customizing pre-trained models for various applications in educational settings. We explain each step in the LLM life cycle and identify potential sources of bias that may arise in the context of education.

Motivation: By defining sources of bias we can add quality checks at specific points in the fine tuning of LLM models for the Dutch context

URL: https://doi.org/10.1111/bjet.13505

TAG: BIAS

5 A conversation around AI observatories

This section presents curated evidence sources across observatories, datasets, data hubs, and periodic reports that monitor AI developments relevant to education. Use the expandable tables below to explore each resource.

5.1 Observatories

Curated portals and platforms that monitor AI policy, standards, incidents, and sectoral developments across health, education, media, urban, and digital domains. Sources include OECD.AI, EU AI Watch, national observatories, and global indexes.

Title details_html
AI incident database
Details

Description: The AI Incident Database is an incident collection initiated by an industrial/non-profit cooperative to enable AI incident avoidance and mitigation. The database supports a variety of research and development use cases with faceted and full text search on more than 2,000 incident reports archived to date.

URL: https://incidentdatabase.ai

AI Observatory Africa
Details

Description: Snapshot of incident database

URL: https://www.africanobservatory.ai/about

Global Index AI
Details

Description: The Global Index is designed to equip governments, civil society, and stakeholders with the evidence needed to advance rights-based principles for the responsible use of AI

URL: https://global-index.ai/

ARTIFICIAL INTELLIGENCE AND DEMOCRATIC VALUES - Library
Details

Description: The < AI & Equality > Toolbox is a collaboration between Women at the Table and EPFL, in consultation with OHCHR.

URL: https://aiequalitytoolbox.com/library/

India AI
Details

Description: Envisioned as an umbrella programme by the Ministry of Electronics and Information Technology (MeitY) for leveraging transformative technologies to foster inclusion, innovation, and adoption for social impact. Consequently, INDIAai (The National AI Portal of India) is poised to assume the pivotal role of a content repository for the INDIAai programme.

URL: https://indiaai.gov.in/

AI in Health Observatory - Tools
Details

Description: The AI in Health Observatory has already identified more than 100 Artificial Intelligence tools in Health from the Catalan Health System (SISCAT), research centres, universities and companies.

URL: https://iasalut.cat/en/observatori-ia-en-salut/registre-dalgorismes/

AI Observatory - Policy
Details

Description: Observatory on AI Policies in Canadian Post-Secondary Education

URL: https://higheredstrategy.com/ai-observatory-home/ai-observatory-policies-and-guidelines/

AI Observatory - Research
Details

Description: Observatory on AI Policies in Canadian Post-Secondary Education

URL: https://higheredstrategy.com/ai-observatory-home/ai-observatory-news-and-research/

AI Standards Hub - Policy
Details

Description: UK - Find AI-related documents published by governmental and other official sources within your area of interest.

URL: https://aistandardshub.org/policy-and-strategy-search/

AI Standards Hub- Standards
Details

Description: Find information on AI-related standards using the search and filtering capabilities below. This database currently covers nearly 300 relevant standards that are being developed or have been published by a range of prominent Standards Development Organisations.

URL: https://aistandardshub.org/ai-standards-search/

AI Standards Hub- Research & Analysis
Details

Description: Find information on AI-related research & analysis using the search and filtering capabilities below. This database currently covers nearly 300 relevant standards that are being developed or have been published by a range of prominent Standards Development Organisations.

URL: https://aistandardshub.org/research-and-analysis-search/

AI Watch - Observatory
Details

Description: Searchable set of currated publications

URL: https://ai-watch.ec.europa.eu/publications_en

AI on Demand - Educational Resources
Details

Description: Welcome to the Education Catalog! Here you can browse, search and access all the AI education resources currently indexed in the AI4EU AI-on-demand platform, including distance courses, onsite courses, and more.

URL: https://www.ai4europe.eu/education/education-catalog

AI4K12 - Learning Materials
Details

Description: The Artificial Intelligence (AI) for K-12 initiative (AI4K12) is jointly sponsored by AAAI and CSTA. The initiative is developing (1) national guidelines for AI education for K-12, (2) an online, curated resource Directory to facilitate AI instruction, and (3) a community of practitioners, researchers, resource and tool developers focused on the AI for K-12 audience. Check out the following information to learn about this initiative.

URL: https://ai4k12.org/resources/list-of-resources/

AI4Media - Resources
Details

Description: he European AI Media Observatory is a knowledge platform that monitors and curates relevant research on AI in media, provides expert perspectives on the potentials and challenges that AI poses for the media sector and allows stakeholders to easily get in touch with relevant experts in the field via our directory.

URL: https://www.ai4media.eu/resources-library/

AI4Media - Feed
Details

Description: he European AI Media Observatory is a knowledge platform that monitors and curates relevant research on AI in media, provides expert perspectives on the potentials and challenges that AI poses for the media sector and allows stakeholders to easily get in touch with relevant experts in the field via our directory.

URL: https://www.ai4media.eu/your-ai-media-feed/

CEDEFOP - Vocational education and training in Europe. Detailed description of VET systems in Europe
Details

Description: This database is the main source of information about vocational education and training (VET) systems in the European Union, Iceland and Norway

URL: https://www.cedefop.europa.eu/en/tools/vet-in-europe/systems/

AI Landscape Startups
Details

Description: It is essential to shine a light on the top AI startups in Europe in order to drive AI adoption and create more partnership opportunities between startups and corporations. By creating a centralized database of quality AI startups, corporations and SMEs will have easier access to AI partners they can trust. With the objective to foster and accelerate AI in Europe, the appliedAI Institute for Europe (Germany), Hub France IA, Ignite Sweden, AI Sweden, RISE Research Institutes of Sweden (Sweden) and the Netherlands AI Coalition (NL AIC) came together in early 2020 to start the initial steps in mapping the AI startup ecosystem in Europe.

URL: https://www.ai-startups-europe.eu/

DESI - Visualization
Details

Description: From 2014 to 2022, the Digital Economy and Society Index (DESI) summarised indicators on Europe’s digital performance and tracked the progress of EU countries.

URL: https://digital-decade-desi.digital-strategy.ec.europa.eu/

European Composite Indicators
Details

Description: European Scorecards (149 - Dec 2023)

URL: https://composite-indicators.jrc.ec.europa.eu/explorer/explorer/indices-and-scoreboards

CEDEFOP - Online Tools
Details

Description: Cedefop is one of the EU’s decentralised agencies. Founded (1) in 1975 and based in Greece since 1995, Cedefop supports development of European vocational education and training (VET) policies and contributes to their implementation. The agency is helping the European Commission, EU Member States and the social partners to develop the right European VET policies. Cedefop: helping develop the right policies to provide the right skills

URL: https://www.cedefop.europa.eu/en/online-tools

Global Observatory of Urban Artificial Intelligence (GOUAI)
Details

Description: The Global Observatory of Urban Artificial Intelligence (GOUAI) is a joint project of CIDOB (Barcelona Centre for International Affairs), and the cities of Barcelona, Amsterdam, and London with the support of UNHabitat. It falls within the frame of the Cities Coalition for Digital Rights, an international city network association that puts together more than 45 cities in the world committed to technological sovereignty and to the promotion of digital rights. GOUAI fills the gap of the lack of consensus and actionable research taking into consideration the local context and aims to contribute to the development of standards through the monitoring of AI initiatives and the research and reflection on key technological issues around AI. The goal is to promote ethical artificial intelligence systems in cities and ensure that algorithmic tools are sustainable, fair, aligned to democratic values, accountable, transparent, cyber secure, and that they safeguard people’s digital rights.

URL: https://gouai.cidob.org/resources/

OECD AI policy Observatory
Details

Description: A global hub for AI policy OECD.AI’s tools, data and other AI policy resources are freely accessible to all actors and stakeholder groups in developed and developing countries.

URL: https://oecd.ai/en/

OECD AI policy Observatory - Education
Details

Description: A global hub for AI policy OECD.AI’s tools, data and other AI policy resources are freely accessible to all actors and stakeholder groups in developed and developing countries.

URL: https://oecd.ai/en/dashboards/policy-areas/PA7

OECD AI policy Observatory - AI & the digital economy
Details

Description: A global hub for AI policy OECD.AI’s tools, data and other AI policy resources are freely accessible to all actors and stakeholder groups in developed and developing countries.

URL: https://oecd.ai/en/dashboards/policy-areas/PA5

OECD AI policy Observatory - AI & Economy
Details

Description: A global hub for AI policy OECD.AI’s tools, data and other AI policy resources are freely accessible to all actors and stakeholder groups in developed and developing countries.

URL: https://oecd.ai/en/dashboards/policy-areas/PA6

OECD programme on AI in Work, Innovation, Productivity and Skills - AI Diffusion
Details

Description: OECD programme on AI in Work, Innovation, Productivity and Skills

URL: https://oecd.ai/en/work-innovation-productivity-skills/key-themes/ai-diffusion

OCEANIS
Details

Description: The Open Community for Ethics in Autonomous and Intelligent Systems (OCEANIS) will provide a high level global forum for discussion, debate and collaboration for organizations interested in the development and use of standards to further the development of autonomous and intelligent systems.

URL: https://ethicsstandards.org/repository/

Map of Science - AI
Details

Description: TO’s Map of Science collects and organizes the world’s research literature, revealing key trends, hotspots, and concepts in global science and technology. The Map of Science includes hundreds of millions of scholarly publications from around the world, algorithmically organized into over 85,000 research clusters. These clusters are groups of papers that cite each other a lot - typically because they share other things in common, such as subject matter, language of publication, or other important characteristics. By filtering clusters according to growth, key topics, countries, and other characteristics, you can quickly pinpoint scientific trends and communities.

URL: https://sciencemap.eto.tech/?all_subjects=artificial+intelligence&mode=summary

DS4Skills - Inventory of Skills and data
Details

Description: The Data Space for Skills (DS4Skills) is a 1-year project aiming to prepare the ground for the development of an open and trusted European Data Space for Skills that supports sharing and accessing skills data.It is funded by the European Commission under the Digital Europe Programme and involves 14 ambitious partners from the industry, education and data ecosystem sectors.

URL: https://inventory.skillsdataspace.eu

5.2 Data sets

Structured datasets covering AI indicators, skills, employment, education, innovation, and digital transformation. Provided by Eurostat, OECD, Stanford HAI, CEDEFOP, and international survey programmes.

Title details_html
AI incident database snapshot
Details

Description: Snapshots of AI incident database

URL: https://incidentdatabase.ai/research/snapshots/

Post-Secondary Employment Outcomes (PSEO)
Details

Description: Post-Secondary Employment Outcomes (PSEO) are experimental tabulations developed by researchers at the U.S. Census Bureau. PSEO data provide earnings and employment outcomes for college and university graduates by degree level, degree major, and post-secondary institution. These statistics are generated by matching university transcript data with a national database of jobs, using state-of-the-art confidentiality protection mechanisms to protect the underlying data.

URL: https://lehd.ces.census.gov/data/pseo_experimental.html

Standford AI Index - data set
Details

Description: The AI Index is an independent initiative at the Stanford Institute for Human-Centered Artificial Intelligence (HAI), led by the AI Index Steering Committee, an interdisciplinary group of experts from across academia and industry. The annual report tracks, collates, distills, and visualizes data relating to artificial intelligence, enabling decision-makers to take meaningful action to advance AI responsibly and ethically with humans in mind.

URL: https://drive.google.com/drive/folders/1ma9WZJzKreS8f2It1rMy_KkkbX6XwDOK

GovTech Maturity Index - Dataset
Details

Description: The GovTech Maturity Index (GTMI) was developed as part of the GovTech Initiative to introduce a measure of GovTech maturity in four focus areas — supporting core government systems, enhancing service delivery, mainstreaming citizen engagement, and fostering GovTech enablers— and to assist practitioners in the design of new digital transformation projects. Constructed for 198 economies, the GTMI is the most comprehensive measure of digital transformation in the public sector. The GTMI is not intended to create a ranking or assess a country’s readiness for or performance of GovTech; rather, it is intended to complement existing tools and diagnostics by providing a baseline and benchmark for GovTech maturity and identifying areas for improvement.

URL: https://datacatalog.worldbank.org/search/dataset/0037889/govtech-dataset

LinkedIn - Dataset
Details

Description: Data that captures the evolution of skill requirements over time across the workforce based on updates to LinkedIn member profiles. This dataset is part of the LinkedIn - World Bank Group partnership, which helps governments and researchers understand rapidly evolving labor markets with detailed and dynamic data. It allows leaders to benchmark and compare labor markets across the world; analyze skills, occupations, migration, and industries; and leverage real-time data to make policy changes. Visualizations for many of these data are available at linkedindata.worldbank.org. The data cover 2015-2019, are refreshed on an annual basis, and are available for 140+ countries.

URL: https://datacatalog.worldbank.org/search/dataset/0038027/Skills—LinkedIn-Data

Digital Intelligence Index (DII) - Dataset
Details

Description: Digital Intelligence Index Digital Intelligence Index (DII) is an interactive research platform of scorecards built to provide evidence-driven, actionable insights on how to enhance digital competitiveness, nurture trust in the digital economy, and foster responsible use of data, AI, and other advanced technologies for enhanced productivity and the greater good. The scorecards and the insights of the DII are meant to inform decision making by technologists, innovators, investors, policymakers, and business leaders on our collective journey from a digital present towards a data-enriched, inclusive, artificial intelligence-augmented future. The current index encompasses the third edition of the Digital Evolution scorecard, following up on our earlier editions (in 2017 and 2014) and the second edition of our Digital Trust scorecard. It paints a picture of global digital development, sheds insight on key factors driving change and momentum, and unpacks the impact digital trust and digital evolution have on a country’s digital competitiveness. This edition of the DII offers data, insights, and international comparisons to guide decision-makers to chart a path out of the pandemic-induced economic challenges of 2020 and toward a data-enabled, artificial intelligence-augmented, and inclusive digital future.

URL: https://digitalplanet.tufts.edu/dii_databooklet/

Global Data Barometer - Dataset
Details

Description: A new global benchmark that tracks the extent to which data is governed, shared, and used for the public good.

URL: https://globaldatabarometer.org/5813cdc8-4f46-4b5e-a626-ea7df3ed2d57

AI & Big Data Global Surveillance Index (2022 updated) - Dataset
Details

Description: his index compiles empirical data on AI and big data surveillance use for 179 countries around the world between 2012 and 2022— the bulk of the sources are found between 2017 and 2022. The index does not distinguish between legitimate and illegitimate uses of AI surveillance techniques. Rather, the purpose of the research is to show how new surveillance capabilities are transforming governments’ ability to monitor and track individuals or groups. Last updated March 2022.

URL: https://data.mendeley.com/datasets/gjhf5y4xjp/4

Secondary use of health data in Western Balkan countries
Details

Description: This report, produced by the ODI and commissioned by Roche, compares the policy readiness of five countries in the Western Balkans region - Bosnia and Herzegovina, Kosovo, Montenegro, North Macedonia and Serbia - for the secondary use of health data. It identifies good practice across the region, policy opportunities and policy needs to improve the European health data policy environment. Report: https://theodi.org/insights/reports/secondary-use-of-health-data-in-western-balkan-countries/

URL: https://drive.google.com/drive/folders/1CvxbjmI9KwI00Uc-8vx1eBsJFDQemEzg

AI Watch - AI TES Dataset 2019
Details

Description: This dataset is the result of a study aimed at mapping the techno-economic segment (TES) of artificial intelligence (AI) from a multidimensional perspective, providing an overview of the worldwide AI landscape in the last decade. The unit of the analysis is the economic agent, or player, that can be a company, university, research institution or governmental authority. The player is expected to have an active role in the segment, with the capability to influence its economic development and future evolution. In this sense, the focus is set on the organisations, and not on individuals, namely the applicant organisation owning the invention in the case of patents, authors’ affiliation in conference proceedings, companies, governmental entities, etc. This dataset targets both industrial and R&D activities. Therefore, players’ economic activities of interest for the analysis of the TES ecosystem include R&D processes (research and innovative developments), general economic processes (industrial production, trade, marketing and other services), firms funding (venture capital funds or other types of investment). The results are provided aggregated by country of location of the institutions concerned.

URL: https://ai-watch.ec.europa.eu/data/ai-tes-dataset-2019_en

AI Watch - Index 2021
Details

Description: The AI Watch Index contains 22 indicators related to the development of AI from several perspectives, organised around 5 dimensions: (i) global view on the AI landscape, (ii) industry, (iii) research and& development (R&D), (iv) technology, and (v) societal aspects. Although the geographical focus is on the EU, when possible the index provides a comparison with major worldwide AI powerhouses, i.e., the US and China, among others. Also, when available, an indicator is provided for the 27 EU Member States.

URL: https://ai-watch.ec.europa.eu/data/ai-watch-index-2021_en

AI Watch - Selected AI Breakthroughs for the AI Watch Timeline
Details

Description: This dataset contains important AI breakthroughs that have been used for visualising the AI history as a timeline on the AI Watch portal

URL: https://ai-watch.ec.europa.eu/data/selected-ai-breakthroughs-ai-watch-timeline_en

CEDEFOP - European Skills Index - Data - 2022
Details

Description: Cedefop’s European Skills Index (ESI) measures the performance of national skills systems using a composite indicator approach. ESI monitors countries’ performance over time and provides insights into possible improvement areas. ESI delivers evidence that supports the EU policy framework for VET and skills, in particular the European Pillar of Social Rights and the European Skills Agenda.

URL: https://www.cedefop.europa.eu/files/esi_2022_scores.xlsx

CEDEFOP - European skills and jobs survey - Dashboard/Publications
Details

Description: The Cedefop European skills and jobs survey (ESJS) is a periodic survey collecting information on the job-skill requirements, digitalisation, skill mismatches and workplace learning of representative samples of European adult workers. A second wave – ESJS2 - carried out in 2021 in all EU Member States, Iceland and Norway, builds on the approach of the first 2014 survey. The ESJS2 aims to inform the policy debate on the impact of digitalisation on the future of work and skills, also in the context of the COVID-19 pandemic. It does so by collecting harmonised international data using a common methodological approach, enabling cross-country comparisons across economic sectors, occupations and other key socioeconomic and demographic variables. Strongly linked to the growth, sustainability and resilience ambitions of the European Skills Agenda and European Digital Strategy, the ESJS2 is a key tool for evidence-based policymaking in VET, skills and employment.

URL: https://www.cedefop.europa.eu/en/tools/european-skills-jobs-survey

Eurostat - Community Innovation Survey
Details

Description: The community innovation survey (CIS) is the reference survey on innovation in enterprises.

URL: https://ec.europa.eu/eurostat/web/microdata/community-innovation-survey

ESCO taxonomy
Details

Description: ESCO can be used by developers as a building block for different types of applications that provide services such as autocomplete, suggestion systems, job search algorithms and job matching algorithms. The ESCO classification is composed of modules that contain elements such as occupations, knowledge, skills and competences, qualifications, and the International Standard Classification of Occupations (ISCO) hierarchy. When combined and interrelated, these modules make up the whole classification. Here you can select your modules and can prepare your custom package of datasets of the ESCO classification to download.

URL: https://esco.ec.europa.eu/en/use-esco/download

European Digital Decade Indicators
Details

Description: Digital Decade and baseline trajectories are presented in the Communication from the Commission establishing Union-level projected trajectories for the digital targets

URL: https://digital-decade-desi.digital-strategy.ec.europa.eu/datasets/dd-trajectories/indicators#digital_decade

Regional innovation scoreboard
Details

Description: The RIS is a regional extension of the European innovation scoreboard (EIS), assessing the innovation performance of European regions on a limited number of indicators

URL: https://research-and-innovation.ec.europa.eu/statistics/performance-indicators/regional-innovation-scoreboard_en

Transitions Performance Index (TPI)
Details

Description: Index that measures the move to prosperous sustainability. The TPI is a scoreboard that monitors and ranks countries based on their 4 transitions to fair and prosperous sustainability. The transition is measured on 4 dimensions economic (education, wealth, labour productivity and research and development intensity, industrial base), social (health life, work and inclusion, free or non-remunerated time, equality), environmental (greenhouse gas emissions reduction, biodiversity, material use, energy productivity), governance (fundamental rights, security, transparency, sound public finances). These measurements are the basis for a new model of prosperity for Europe and the world. A model which focuses on resilience, inclusiveness and sustainability and which supports the EU’s 2022 Annual Sustainable Growth Strategy.

URL: https://research-and-innovation.ec.europa.eu/strategy/support-policy-making/support-national-research-and-innovation-policy-making/transitions-performance-index-tpi_en

Eurofound Surveys - Skills
Details

Description: Eurofound has developed several regularly repeated pan-European surveys to contribute to the planning and establishment of better living and working conditions. The surveys offer a unique source of comparative information on the quality of living and working conditions across the EU.

URL: https://www.eurofound.europa.eu/en/data-catalogue?f%5B0%5D=ef_data_catalogue_text%3Askill&page=1

OECD Skills Survey
Details

Description: The Programme for the International Assessment of Adult Competencies (PIAAC) is a programme of assessment and analysis of adult skills. The major survey conducted as part of PIAAC is the Survey of Adult Skills. The Survey measures adults’ proficiency in key information-processing skills - literacy, numeracy and problem solving - and gathers information and data on how adults use their skills at home, at work and in the wider community. This international survey is conducted in over 40 countries/economies and measures the key cognitive and workplace skills needed for individuals to participate in society and for economies to prosper.

URL: https://www.oecd.org/skills/piaac/data/

Euro barometer - Polling - SP532 : The digital decade
Details

Description: Eurobarometer is the polling instrument used by the European Commission

URL: https://data.europa.eu/data/datasets/s2959_99_1_sp532_eng?locale=en

Euro barometer - Polling - FL496 : SMEs and Cybercrime
Details

Description: Eurobarometer is the polling instrument used by the European Commission

URL: https://data.europa.eu/data/datasets/s2280_fl496_eng?locale=en

Euro barometer - Polling - 503: Attitudes towards the impact of digitalisation on daily lives
Details

Description: Eurobarometer is the polling instrument used by the European Commission

URL: https://data.europa.eu/data/datasets/s2228_92_4_503_eng?locale=en

Google - Trends
Details

Description: Trends around the use of words have been used as independent variables in Economic models around AI

URL: https://trends.google.com/trends/

Artificial Intelligence patent dataset
Details

Description: Measuring Patient activity around AI as early warning has been reported in numerous research papers

URL: https://www.uspto.gov/ip-policy/economic-research/research-datasets/artificial-intelligence-patent-dataset

GENDER EQUALITY STRATEGY Monitoring Portal
Details

Description: Gender Equality Monitor including datasets

URL: https://composite-indicators.jrc.ec.europa.eu/ges-monitor/resources

European Composite Indicators - digital-nomad-index
Details

Description: The Digital Nomad Index measures how countries are suited to be new digital working homes. The index looks at a range of factors from the cost of internet speed and the Global Happiness Index to the monthly rent of a one-bedroom apartment and the number of searches for remote jobs among others.

URL: https://composite-indicators.jrc.ec.europa.eu/explorer/explorer/indices/dni/digital-nomad-index

European Composite Indicators - Cyber Defense index
Details

Description: The Cyber Defense Index (CDI) measures the extent to which the world’s 20 largest and most digitally forward economies have adopted technology and digital practices to resist cyberattacks, and how well their governments and policy frameworks promote cybersecure digital transactions.

URL: https://composite-indicators.jrc.ec.europa.eu/explorer/explorer/indices/cdfi/cyber-defense-index

European Composite Indicators - European Index of Digital Entrepreneurship Systems
Details

Description: The EIDES attempts to measure both physical and digital conditions for stand-up, start-up and scale-up ventures in EU countries by building on the existing JRC Entrepreneurship and Scale-up Indices (ESIS). The EIDES encompasses eight pillars: four pillars that represent the General Framework Conditions, and four pillars that represent the Systemic Framework Conditions

URL: https://composite-indicators.jrc.ec.europa.eu/explorer/explorer/indices/eides/european-index-of-digital-entrepreneurship-systems

European Composite Indicators - ICT Regulatory Tracker
Details

Description: The ICT Regulatory Tracker (ICTRT) pinpoints the changes taking place in the ICT regulatory environment. It facilitates benchmarking and the identification of trends in ICT legal and regulatory frameworks in 193 countries. The ICTRT does not measure the quality, the level of implementation or the performance of regulatory frameworks in place, but records their existence and features.

URL: https://composite-indicators.jrc.ec.europa.eu/explorer/explorer/indices/ictrt/ict-regulatory-tracker

European Composite Indicators - Global Knowledge Index
Details

Description: The Global Knowledge Index (GKI) represents an important addition to the global repository of knowledge on development, providing updated, comprehensive, and reliable data that better enable countries and decision makers to understand and respond more clearly to the transformations and challenges related to knowledge-based development.

URL: https://composite-indicators.jrc.ec.europa.eu/explorer/explorer/indices/gki/global-knowledge-index

European Composite Indicators - Global Sustainable Competitiveness Index
Details

Description: The Global Sustainable Competitiveness Index measures national development and green growth since 2012, based on quantitative indicators derived from recognized international organisations to present a comprehensive view of a nation’s performance and future potential.

URL: https://composite-indicators.jrc.ec.europa.eu/explorer/explorer/indices/gsci/global-sustainable-competitiveness-index

European Composite Indicators - Global Data Barometer
Details

Description: The Global Data Barometer is a tool that seeks to measure the state of data around the world. The Barometer assesses countries around the world on a set of robust comparative metrics focused on data for public good.

URL: https://composite-indicators.jrc.ec.europa.eu/explorer/explorer/indices/gdb/global-data-barometer

European Composite Indicators - Green, Digital and Competitive SME Index
Details

Description: The Green, Digital and Competitive SME Index is a multi-year project intended to support SMEs in the twin transition and accompany the European Commission’s European Green Deal programme. The Index is built around an innovative methodology with three pillars: 1) Green: How fast are SMEs moving to lower emissions and increase energy efficiency? 2) Digital: How quickly are those enterprises embracing modern technology across their organisations? And 3) Competitive: How quick are they growing and how well performing outside of home markets?

URL: https://composite-indicators.jrc.ec.europa.eu/explorer/explorer/indices/gdcsi/green-digital-and-competitive-sme-index

European Composite Indicators - Inclusive Internet Index
Details

Description: The index provides a rigorous benchmark of national-level Internet inclusion across four categories: Availability, Affordability, Relevance and Readiness. The goal is to give policymakers and influencers a clearer understanding of the factors that contribute to wide and sustainable inclusion and to provide a comparison across countries of the enabling environment for adoption and productive use of the Internet.

URL: https://composite-indicators.jrc.ec.europa.eu/explorer/explorer/indices/3i-index/inclusive-internet-index

European Composite Indicators - Innovation Output Indicator
Details

Description: The Innovation Output Indicator offers an output-oriented measure of innovation performance of EU Countries. It measures countries’ capacity to derive economic benefits from innovation.

URL: https://composite-indicators.jrc.ec.europa.eu/explorer/explorer/indices/ioi/innovation-output-indicator

European Composite Indicators - Mobility Scoreboard in Initial Vocational Education and Training
Details

Description: The European Commission’s Mobility Scoreboard follows up on the 2011 ‘Youth on the Move’ Recommendation of the Council of the European Union, providing a framework for monitoring progress made by European countries in creating a positive environment supporting learner mobility. The Mobility Scoreboard in Initial Vocational Education and Training (IVET) is a tool for monitoring developments in IVET mobility policies.

URL: https://composite-indicators.jrc.ec.europa.eu/explorer/explorer/scoreboards/ms-ivec/mobility-scoreboard-in-initial-vocational-education-and-training

European Composite Indicators - Open Data Maturity
Details

Description: The Open Data Maturity study serves as a benchmark to gain insights into the development achieved in the field of open data in Europe. It assesses the level of maturity against four dimensions: policy, portal, impact, and quality.

URL: https://composite-indicators.jrc.ec.europa.eu/explorer/explorer/indices/odm/open-data-maturity

European Composite Indicators - Open-Useful-Reusable Government data Index
Details

Description: The OECD Open-Useful-Re-Usable Government Data Index (OURdata Index) presents information on open government data (OGD) policies and practices in OECD countries based on the International Open Data Charter. Three pillars form the overall index (data availability, data accessibility, and government support for the re-use of data).

URL: https://composite-indicators.jrc.ec.europa.eu/explorer/explorer/indices/ourdata/open-useful-reusable-government-data-index

European Composite Indicators - Mobility Scoreboard in Higher Education
Details

Description: The Higher Education (HE) Mobility Scoreboard is a tool for monitoring developments in HE mobility policies and the implementation of the 2011 Council recommendation ‘Youth on the move’ in European countries. The scoreboard is a mechanism for assisting policy-making in international learning mobility.

URL: https://composite-indicators.jrc.ec.europa.eu/explorer/explorer/scoreboards/ms-he/mobility-scoreboard-in-higher-education

European Composite Indicators - SME Performance Review
Details

Description: The SME Performance Review is one of the main tools the European Commission to monitor and assess countries’ progress in implementing the SME strategy and the Small Business Act (SBA), mapping SME performance in EU countries and partner countries.

URL: https://composite-indicators.jrc.ec.europa.eu/explorer/explorer/scoreboards/smepr/sme-performance-review

European Composite Indicators - Women in Digital Scoreboard
Details

Description: The Women in Digital Scoreboard (WiD scoreboard) is one of the actions put in place to assess women’s inclusion in digital jobs, careers and entrepreneurship. The WiD scoreboard assesses Member States’ performance in the areas of Internet use, Internet user skills as well as specialist skills and employment based on 12 indicators.

URL: https://composite-indicators.jrc.ec.europa.eu/explorer/explorer/indices/wids/women-in-digital-scoreboard

CORDIS - EU research projects under HORIZON EUROPE
Details

Description: This dataset contains information about projects and their results funded by the European Union under the Horizon Europe framework programme for research and innovation from 2021 to 2027. Reference data (programmes, topics, topic keywords funding schemes (types of action), organisation types and countries) can be found in this dataset: https://data.europa.eu/euodp/en/data/dataset/cordisref-data

URL: https://data.europa.eu/data/datasets/cordis-eu-research-projects-under-horizon-europe-2021-2027?locale=en

AIAAIC Repository
Details

Description: The independent, open, public interest resource detailing incidents and controversies driven by and relating to artificial intelligence, algorithms, and automation.

URL: https://docs.google.com/spreadsheets/d/1Bn55B4xz21-_Rgdr8BBb2lt0n_4rzLGxFADMlVW0PYI/edit#gid=888071280

Mitre
Details

Description: MITRE ATLAS™ (Adversarial Threat Landscape for Artificial-Intelligence Systems) is a globally accessible, living knowledge base of adversary tactics and techniques based on real-world attack observations and realistic demonstrations from AI red teams and security groups. There are a growing number of vulnerabilities in AI-enabled systems, as the incorporation of AI increases the attack surface of existing systems beyond those of traditional cyber-attacks. We developed ATLAS to raise awareness of these unique and evolving vulnerabilities, as the global community starts to incorporate AI into more systems. ATLAS is modeled after the MITRE ATT&CK® framework and its tactics, techniques, and procedures (TTPs) are complementary to those in ATT&CK.

URL: https://atlas.mitre.org/

Biases in AI example tracker
Details

Description: Google Doc

URL: https://docs.google.com/spreadsheets/d/1eyZZW7eZAfzlUMD8kSU30IPwshHS4ZBOyZXfEBiZum4/edit#gid=910445724

OECD programme on AI in Work, Innovation, Productivity and Skills - Jobs and Skills
Details

Description: OECD programme on AI in Work, Innovation, Productivity and Skills

URL: https://oecd.ai/en/data?selectedArea=ai-jobs-and-skills

OECD AI Incidents Monitor (AIM)
Details

Description: The OECD AI Incidents Monitor (AIM) documents AI incidents to help policymakers, AI practitioners, and all stakeholders worldwide gain valuable insights into the incidents and hazards that concretise AI risks. Over time, AIM will help to show patterns and establish a collective understanding of AI incidents and their multifaceted nature and serve as an important tool for trustworthy AI.

URL: https://oecd.ai/en/incidents?

ONS (UK) AI datasets/Publications - Awareness
Details

Description: UK’s largest independent producer of official statistics and its recognised national statistical institute. We are responsible for collecting and publishing statistics related to the economy, population and society at national, regional and local levels. We also conduct the census in England and Wales every 10 years.

URL: https://www.ons.gov.uk/businessindustryandtrade/itandinternetindustry/datasets/artificialintelligenceaiawarenessuseandimpactgreatbritain

Onet Data files
Details

Description: The ONET database contains a rich set of variables that describe work and worker characteristics, including skill requirements. Developers and other customers are encouraged to incorporate the ONET database within their products, services, and research. This section contains an overview of the most recent database and a variety of database download options, as well as format-specific data dictionaries.

URL: https://www.onetcenter.org/database.html#all-files

Artificial intelligence (AI) in Sweden 2019
Details

Description: Artificial intelligence (AI) in Sweden 2019 - Dataset and publication

URL: https://www.scb.se/en/finding-statistics/statistics-by-subject-area/education-and-research/research/research-and-development-in-sweden/pong/publications/artificial-intelligence-ai-in-sweden-2019/

Impact of Artificial Intelligence on Jobs
Details

Description: This article is a analysis on how the use of AI and automation will lead to disturbance in society and cause large scale unemployment. (vellikad2019?)

URL: https://data.mendeley.com/datasets/6x3ms7857b/1

Observatory of Algorithms with Social Impact
Details

Description: The Observatory of Algorithms with Social Impact aims to clarify the basic concepts as definitions of algorithm, machine learning and other relevant concepts. But OASI also contains a Register of algorithms being developed and implemented by different institutions and organisations across the world. The OASI Register gathers, classifies and makes algorithms searchable, and is regularly updated with new content.

URL: https://eticas.tech/oasi

Cedefop - Skills Indicators
Details

Description: Indicators are the building blocks of skills intelligence visualisations. They are based on datasets (such as Labour Force Survey, or Cedefop Skills Forecast), sourced by various organisations (such as Eurostat or Cedefop). An indicator is a slice of the data, providing a particular piece of information, such as future employment growth, unemployment rate or task importance. This information can usually be broken down even further - most often by occupation, education, age, gender or sector. Indicators can also relate to most important topics of Cedefop work. Use the filter options to find out which indicators relate to particular sources, datasets or topics.

URL: https://www.cedefop.europa.eu/en/tools/skills-intelligence/indicators

EU Legislation Digitalization
Details

Description: This Bruegel dataset seeks to provide a comprehensive view of:

European Union laws relevant to digitalisation enacted previously, measures that have been enacted or might still be enacted during the current legislative session (2019 - 2024), and ongoing EU policy initiatives that might lead to new legislation in the foreseeable future; For each of the years 2020 - 2023, a figure showing the actual or expected progress through the EU legislative process of the five measures introduced that year that the authors consider to be the most significant;

URL: https://www.bruegel.org/dataset/dataset-eu-legislation-digital-world

OECD - PIAC dataset
Details

Description: Twenty-four countries participated in Round 1 of the Survey of Adult Skills (PIAAC), with data collection taking place from 1 August 2011 to 31 March 2012 in most countries. Nine countries took part in Round 2 of the assessment, with data collection taking place from April 2014 to end-March 2015. Six countries participated in Round 3, with data collection taking place from July to December 2017.

URL: https://www.oecd.org/skills/piaac/publicdataandanalysis/

European enterprise survey on the use of technologies based on artificial intelligence
Details

Description: This study was carried out by Ipsos for the European Commission. It has found that 42% of enterprises currently use at least one AI technology, a quarter of them use at least two types, and 18% have plans to adopt AI technologies in the next two years. Three key internal barriers to AI adoption are difficulties in hiring new staff with the right skills (57%), the cost of adoption (52%) and the cost of adapting operational processes (49%).

URL: https://digital-strategy.ec.europa.eu/en/library/european-enterprise-survey-use-technologies-based-artificial-intelligence

EFF surveillance database
Details

Description: Surveilance Database

URL: https://atlasofsurveillance.org/search

TIOBE index
Details

Description: The TIOBE Programming Community index is an indicator of the popularity of programming languages. The index is updated once a month. The ratings are based on the number of skilled engineers world-wide, courses and third party vendors. Popular search engines such as Google, Bing, Yahoo!, Wikipedia, Amazon, YouTube and Baidu are used to calculate the ratings. It is important to note that the TIOBE index is not about the best programming language or the language in which most lines of code have been written.

URL: https://www.tiobe.com/tiobe-index/

Stack overflow developer survey of tooling - AI/ML dashboard
Details

Description: EC project databases

URL: https://survey.stackoverflow.co/2023/#section-most-popular-technologies-asynchronous-tools

Common crawl
Details

Description: The Common Crawl (CC) is a corpus containing over 5 billion (5.000.000.000!) websites in raw format (WARC), meta (WAT) data and text (WET). This takes up about a petabyte of storage or maybe some more terabytes. Doesn’t matter at this volume size. This amount of websites in multiple languages has been acquired in the last 7 years and is distributed with a CC. In April alone they crawled about 80 TB.

URL: https://blog.entropic-data.com/2018/05/08/defining-data-science-using-the-common-crawl-web-corpus-1/

EUrostat - Digitialisation in Europe 2023 edition
Details

Description: This publication provides easy to understand statistics on information and communication technologies (ICT) and the way people and businesses use digital technologies.

URL: https://ec.europa.eu/eurostat/web/interactive-publications/digitalisation-2023#digital-transformation

5.3 Data Hubs

Aggregated platforms and queryable databases from Eurostat, ITU, ILO, OECD, and national statistical offices, offering multi-domain data on labour, education, innovation, law, and demographics.

Title details_html
Longitudinal Employer-Household Dynamics
Details

Description: LEHD makes available several data products that may be used to research and characterize workforce dynamics for specific groups. These data products include online applications, public-use data, and restricted-use microdata. The Quarterly Workforce Indicators (QWI), LEHD Origin-Destination Employment Statistics (LODES), Job-to-Job Flows (J2J), and Post-Secondary Employment Outcomes (PSEO) are available online for public use. Confidential microdata are available to qualified researchers with approved projects through restricted access use in Federal Statistical Research Data Centers (FSRDCs).

URL: https://lehd.ces.census.gov/data/#lodes

ITU datahub
Details

Description: The world’s richest source of ICT statistics and regulatory information Track the digital transformation with the ITU DataHub, featuring hundreds of ICT indicators on connectivity, markets, affordability, trust governance, and sustainability. Find, compare, and download data for nearly 200 economies.

URL: https://datahub.itu.int/indicators/

BTI transformation index
Details

Description: The BTI is the result of the collaboration of nearly 300 country and regional experts from leading universities and think tanks worldwide. The project analyzes and compares transformation processes towards democracy and inclusive market economy worldwide. The BTI aims to identify successful strategies for steering change.

URL: https://bti-project.org/en/downloads

Bureau of Labor Statistics - Dathub
Details

Description: In 1884, the Bureau of Labor was established in the Department of Interior. In 1888, it became an independent department for nearly 15 years before being incorporated into the Department of Commerce and Labor in 1903. In 1913, the Bureau of Labor Statistics (BLS) was transferred to the newly created Department of Labor.

URL: https://www.bls.gov/data/

Eurostat - Datahub
Details

Description: Eurostat is the statistical office of the European Union.

URL: https://ec.europa.eu/eurostat/web/main/data/database

Eurostat - Censushub
Details

Description: Eurostat is the statistical office of the European Union.

URL: https://ec.europa.eu/CensusHub2/selectHyperCube?qhc=false#

Eramus plus - Search Engine
Details

Description: rasmus+ is the EU’s programme to support education, training, youth and sport in Europe. It has an estimated budget of €26.2 billion. This is nearly double the funding compared to its predecessor programme (2014-2020). The 2021-2027 programme places a strong focus on social inclusion, the green and digital transitions, and promoting young people’s participation in democratic life

URL: https://erasmus-plus.ec.europa.eu/projects/search

Eurofound Survey
Details

Description: Eurofound has developed several regularly repeated pan-European surveys to contribute to the planning and establishment of better living and working conditions. The surveys offer a unique source of comparative information on the quality of living and working conditions across the EU.

URL: https://www.eurofound.europa.eu/en/data

EU Internet Standards Deployment Monitoring Website
Details

Description: The dashboards visualise the deployment indicators of 18 key Internet standards that help to secure the Internet and support its constant technological evolution.

URL: https://ec.europa.eu/internet-standards/downloads.html

Euro barometer - Pololing
Details

Description: Eurobarometer is the polling instrument used by the European Commission, the European Parliament and other EU institutions and agencies to monitor regularly the state of public opinion in Europe on issues related to the European Union as well as attitudes on subjects of political or social nature. Eurobarometer provides quality and relevant data for experts in public opinion, researchers, media and the public.

URL: https://europa.eu/eurobarometer/surveys/browse/all

Espacenet - Patient Hub
Details

Description: Measuring Patient activity around AI as early warning has been reported in numerous research papers

URL: https://worldwide.espacenet.com/

Google - Patient Hub
Details

Description: Measuring Patient activity around AI as early warning has been reported in numerous research papers

URL: https://patents.google.com/

Emerging Technology Observatory
Details

Description: The Emerging Technology Observatory creates free, high-quality data resources to inform critical decisions on emerging technology issues.

URL: https://eto.tech/tools/

Erasmus Plus Projects associated with AI
Details

Description: EU programme for education, training, youth and sport. Search Engine

URL: https://erasmus-plus.ec.europa.eu/projects/search/?page=1&sort=&domain=eplus2021&view=list&map=false&keyword=AI&searchType=projects

EuroLex
Details

Description: Access to European Law

URL: https://eur-lex.europa.eu/homepage.html

International Labour Organization (ILO)
Details

Description: The only tripartite U.N. agency, since 1919 the ILO brings together governments, employers and workers of 187 Member States , to set labour standards, develop policies and devise programmes promoting decent work for all women and men.Labour statistics play an essential role in the efforts of member States to achieve decent work for all and for the ILO’s support of these efforts. These statistics are needed for the development and evaluation of policies towards this goal and for assessing progress towards decent work. They are also an important tool for information and analysis, helping to increase understanding of common problems, explain actions and mobilize interest.

URL: https://www.ilo.org/global/statistics-and-databases/lang–en/index.htm

OECD programme on AI in Work, Innovation, Productivity and Skills - Live Data
Details

Description: OECD programme on AI in Work, Innovation, Productivity and Skills

URL: https://oecd.ai/en/work-innovation-productivity-skills/live-data

ONS (UK) AI datasets/Publications
Details

Description: UK’s largest independent producer of official statistics and its recognised national statistical institute. We are responsible for collecting and publishing statistics related to the economy, population and society at national, regional and local levels. We also conduct the census in England and Wales every 10 years.

URL: https://www.ons.gov.uk/search?q=Artificial

World Inequality Database on Education
Details

Description: The World Inequality Database on Education (WIDE) highlights the powerful influence of circumstances, such as wealth, gender, ethnicity and location, over which people have little control but which play an important role in shaping their opportunities for education and life. It draws attention to unacceptable levels of education inequality across countries and between groups within countries, with the aim of helping to inform policy design and public debate.

URL: https://www.education-inequalities.org

UNESCO Institute for Statistics (UIS)
Details

Description: The UNESCO Institute for Statistics (UIS) is the official and trusted source of internationally-comparable data on education, science, culture and communication. As the official statistical agency of UNESCO, the UIS produces a wide range of state-of-the-art databases to fuel the policies and investments needed to transform lives and propel the world towards its development goals. The UIS provides free access to data for all UNESCO countries and regional groupings from 1970 to the most recent year available. The UIS encourages developers and researchers to build websites and applications that make rich use of UIS dissemination data. In addition to a powerful standards-based API, the UIS supports a data browser and a bulk data download service (BDDS).

URL: http://data.uis.unesco.org/

ITU datahub
Details

Description: The world’s richest source of ICT statistics and regulatory information Track the digital transformation with the ITU DataHub, featuring hundreds of ICT indicators on connectivity, markets, affordability, trust governance, and sustainability. Find, compare, and download data for nearly 200 economies.

URL: https://datahub.itu.int/indicators/

EC project databases
Details

Description: EC project databases

URL: https://research-and-innovation.ec.europa.eu/projects/project-databases_en

Stack overflow developer survey of tooling - Yearly dataset
Details

Description: EC project databases

URL: https://insights.stackoverflow.com/survey

5.4 Reporting

Periodic index reports and recurring assessments tracking AI readiness, innovation, education, employment, and democratic values. Includes Stanford AI Index, Global Innovation Index, GEM Report, and EU scoreboards.

Title details_html
Standford AI Index
Details

Description: The AI Index is an independent initiative at the Stanford Institute for Human-Centered Artificial Intelligence (HAI), led by the AI Index Steering Committee, an interdisciplinary group of experts from across academia and industry. The annual report tracks, collates, distills, and visualizes data relating to artificial intelligence, enabling decision-makers to take meaningful action to advance AI responsibly and ethically with humans in mind.

URL: https://aiindex.stanford.edu/report/

Global Innovation Index
Details

Description: The Global Innovation Index (GII) takes the pulse of innovation against a background of an economic and geopolitical environment fraught with uncertainty. It reveals the most innovative economies in the world, ranking the innovation performance of around 132 economies while highlighting innovation strengths and weaknesses.

URL: https://www.wipo.int/global_innovation_index/en/

Global Innovation Index - Dashboard
Details

Description: The Global Innovation Index (GII) takes the pulse of innovation against a background of an economic and geopolitical environment fraught with uncertainty. It reveals the most innovative economies in the world, ranking the innovation performance of around 132 economies while highlighting innovation strengths and weaknesses.

URL: https://www.wipo.int/gii-ranking/en/

UN E-Government Survey 2022
Details

Description: The United Nations E-Government Survey 2022 is the 12th edition of the United Nations’ assessment of the digital government landscape across all 193 Member States. The E-Government Survey is informed by over two decades of longitudinal research, with a ranking of countries based on the United Nations E-Government Development Index (EGDI), a combination of primary data (collected and owned by the United Nations Department of Economic and Social Affairs) and secondary data from other UN agencies.

This edition of the Survey includes data analysis in global and regional contexts, a study of local e-government development based on the United Nations Local Online Service Index (LOSI), consideration of inclusion in the hybrid digital society, and a concluding chapter that outlines the trends and developments related to the future of digital government. As wish all editions, it features extensive annexes on its data, methodology and related pilot study initiatives.

URL: https://publicadministration.un.org/egovkb/en-us/Reports/UN-E-Government-Survey-2022

GovTech Maturity Index
Details

Description: The GovTech Maturity Index (GTMI) was developed as part of the GovTech Initiative to introduce a measure of GovTech maturity in four focus areas — supporting core government systems, enhancing service delivery, mainstreaming citizen engagement, and fostering GovTech enablers— and to assist practitioners in the design of new digital transformation projects. Constructed for 198 economies, the GTMI is the most comprehensive measure of digital transformation in the public sector. The GTMI is not intended to create a ranking or assess a country’s readiness for or performance of GovTech; rather, it is intended to complement existing tools and diagnostics by providing a baseline and benchmark for GovTech maturity and identifying areas for improvement.

URL: https://www.worldbank.org/en/programs/govtech/gtmi

Digital Intelligence Index (DII)
Details

Description: Digital Intelligence Index Digital Intelligence Index (DII) is an interactive research platform of scorecards built to provide evidence-driven, actionable insights on how to enhance digital competitiveness, nurture trust in the digital economy, and foster responsible use of data, AI, and other advanced technologies for enhanced productivity and the greater good. The scorecards and the insights of the DII are meant to inform decision making by technologists, innovators, investors, policymakers, and business leaders on our collective journey from a digital present towards a data-enriched, inclusive, artificial intelligence-augmented future. The current index encompasses the third edition of the Digital Evolution scorecard, following up on our earlier editions (in 2017 and 2014) and the second edition of our Digital Trust scorecard. It paints a picture of global digital development, sheds insight on key factors driving change and momentum, and unpacks the impact digital trust and digital evolution have on a country’s digital competitiveness. This edition of the DII offers data, insights, and international comparisons to guide decision-makers to chart a path out of the pandemic-induced economic challenges of 2020 and toward a data-enabled, artificial intelligence-augmented, and inclusive digital future.

URL: https://digitalplanet.tufts.edu/digitalintelligence/

ARTIFICIAL INTELLIGENCE AND DEMOCRATIC VALUES - Report
Details

Description: n 2020, the Center for AI and Digital Policy published the first worldwide assessment of AI policies and practices. Artificial Intelligence and Democratic Values rated and ranked 30 countries, based on a rigorous methodology and 12 metrics established to assess alignment with democratic values. The 2021 AI Index expands the global coverage from 30 countries to 50 countries, acknowledges the significance of the UNESCO Recommendation on AI ethics, and reviews earlier country ratings. The 2021 report is the result of the work of more than 100 AI policy experts in almost 40 countries.The 2022 AI Index covers 75 countries. The metrics have been updated to take into consideration the implementation of the 2021 UNESCO Recommendation on AI Ethics as well as the 2022 Global Privacy Assembly Resolution on Facial Recognition. Country reports headings have been harmonized to ease reference and comparative analysis. Relevant regional frameworks have been systematically taken into account in the country reports.The 2022 report is the result of the work of more than 200 AI policy experts in almost 60 countries.

URL: https://www.caidp.org/reports/aidv-2022/

AI Readiness Index
Details

Description: Oxford Insights remains committed to providing valuable insights at the intersection of government and AI. This year we assess the AI readiness of 193 governments across the world. We are also introducing an interactive map to make our data more accessible!

URL: https://oxfordinsights.com/ai-readiness/ai-readiness-index/

WIPO The Future of Jobs Report
Details

Description: The Future of Jobs Report 2023 explores how jobs and skills will evolve over the next five years. This fourth edition of the series continues the analysis of employer expectations to provide new insights on how socio-economic and technology trends will shape the workplace of the future.

URL: https://www.weforum.org/publications/the-future-of-jobs-report-2023/

GLOBAL EDUCATION MONITORING REPORT
Details

Description: (globale2023?) Global education monitoring report, 2023: technology in education: a tool on whose terms?

URL: https://unesdoc.unesco.org/ark:/48223/pf0000385723

OECD Employment Outlook
Details

Description: OECD’s annual report on jobs and employment in OECD countries. Each edition reviews recent trends, policy developments, and prospects. A statistical annex provides data on unemployment rates, incidence of part-time employment, employment/population ratios, and activity rates. Also included are data on expenditure on labour market programmes, average annual wages, and earnings dispersion. Special Chapters examine issues of topical interest.

URL: https://www.oecd-ilibrary.org/employment/oecd-employment-outlook-2023_08785bba-en

6 Literature to help refine the search term

This table lists candidate citations for the systematic review. Use the Priority and Reviewer columns to track your screening assessment of each paper.

Citations Priority_And_Notes Reviewer
(Lee, Lee, et al. 2026)
(Lee, Ahn, et al. 2026)
(Lin et al. 2026)
(Wei et al. 2025)
(Quttainah et al. 2024)
(Nikiforova-Ilieva and Georgiev 2025)
(Shahzad et al. 2025)
(Lee et al. 2024)
(Lelièvre et al. 2025)
(Brown 2024)
(Alzahrani et al. 2024)
(Jiang et al. 2025)
(Liu et al. 2024)
(Rooein et al. 2024)
(Jeon et al. 2026)
(Zhou et al. 2025)
(Balepur et al. 2026)
(Moëll et al. 2025)
(Qian et al. 2026)
(Jiao et al. 2025)
(Hamna et al. 2026)
(Li et al. 2025)
(Singh et al. 2025)
(Chen et al. 2025)
(Zhang et al. 2025)
(Arora et al. 2025)
(Lunardi et al. 2025)
(Potamitis et al. 2025)
(Wang et al. 2025)
(Shashidhar et al. 2025)
(Feuer et al. 2025)
(Banyas et al. 2025)
(Mohammadi et al. 2025)
(Zhu et al. 2026)
(Kim et al. 2025)
(Vries et al. 2023)
(Vanroy 2023)
(Vanroy 2024)
(Noels et al. 2024)
(Thellmann et al. 2024)
(Barth and Rehm 2025)
(Vintar et al. 2025)
(Martins et al. 2024)
(Son et al. 2024)
(Chang et al. 2024)
(Ghosh et al. 2024)
(Gao et al. 2025)
(Jegham et al. 2025)
(Mehditabar et al. 2025)
(Yuan et al. 2025)
(Husom et al. 2025)
(Wu et al. 2025)
(Dauner and Socher 2025)
(Vijay and Vadapalli 2025)
(Huang et al. 2025)
(Budler et al. 2025)
(Cámara et al. 2024)
(McIntosh et al. 2026)
(Guo et al. 2023)
(Xu et al. 2024)
(Shool et al. 2025)
(Banerjee et al. 2024)
(Polo et al. 2024)
(Fine Licht 2023)
(Pomerenke et al. 2025)

7 Acknowledgements

Assisted Authoring

This document was created with the assistance of opencode, an AI-powered coding assistant, running DeepSeek V4 Flash as its underlying language model. The tool was used for drafting, editing, and formatting throughout the preparation of this report.

References

Alzahrani, Norah, Hisham Abdullah Alyahya, Yazeed Alnumay, et al. 2024. When Benchmarks Are Targets: Revealing the Sensitivity of Large Language Model Leaderboards. https://doi.org/10.48550/ARXIV.2402.01781.
Arora, Rahul K., Jason Wei, Rebecca Soskin Hicks, et al. 2025. HealthBench: Evaluating Large Language Models Towards Improved Human Health. https://doi.org/10.48550/ARXIV.2505.08775.
Balepur, Nishant, Bhavya Rajasekaran, Jane Oh, et al. 2026. BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice Benchmarks. https://doi.org/10.48550/ARXIV.2602.06221.
Banerjee, Sourav, Ayushi Agarwal, and Eishkaran Singh. 2024. The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance? https://doi.org/10.48550/ARXIV.2412.03597.
Banyas, Peter, Shristi Sharma, Alistair Simmons, and Atharva Vispute. 2025. ConsistencyAI: A Benchmark to Assess LLMs’ Factual Consistency When Responding to Different Demographic Groups. https://doi.org/10.48550/ARXIV.2510.13852.
Barth, Fabio, and Georg Rehm. 2025. Multilingual European Language Models: Benchmarking Approaches and Challenges. https://doi.org/10.48550/ARXIV.2502.12895.
Brown, Nik Bear. 2024. Enhancing Trust in LLMs: Algorithms for Comparing and Interpreting LLMs. https://doi.org/10.48550/ARXIV.2406.01943.
Budler, Leona Cilar, Hongyu Chen, Aokun Chen, et al. 2025. “A Brief Review on Benchmarking for Large Language Models Evaluation in Healthcare.” WIREs Data Mining and Knowledge Discovery 15 (2). https://doi.org/10.1002/widm.70010.
Cámara, Javier, Lola Burgueño, and Javier Troya. 2024. “Towards Standardized Benchmarks of LLMs in Software Modeling Tasks: A Conceptual Framework.” Software and Systems Modeling 23 (6): 1309–18. https://doi.org/10.1007/s10270-024-01206-9.
Chang, Yupeng, Xu Wang, Jindong Wang, et al. 2024. “A Survey on Evaluation of Large Language Models.” ACM Transactions on Intelligent Systems and Technology 15 (3): 1–45. https://doi.org/10.1145/3641289.
Chen, Zhijun, Xiaodong Lu, Jingzheng Li, et al. 2025. Harnessing Multiple Large Language Models: A Survey on LLM Ensemble. https://doi.org/10.48550/ARXIV.2502.18036.
Dauner, Maximilian, and Gudrun Socher. 2025. “Energy Costs of Communicating with AI.” Frontiers in Communication 10 (June). https://doi.org/10.3389/fcomm.2025.1572947.
Feuer, Benjamin, Chiung-Yi Tseng, Astitwa Sarthak Lathe, Oussama Elachqar, and John P Dickerson. 2025. When Judgment Becomes Noise: How Design Failures in LLM Judge Benchmarks Silently Undermine Validity. https://doi.org/10.48550/ARXIV.2509.20293.
Fine Licht, Karl de. 2023. “Integrating Large Language Models into Higher Education: Guidelines for Effective Implementation.” IS4SI Summit 2023, August 11, 65. https://doi.org/10.3390/cmsf2023008065.
Gao, Yan, Massimo Roberto Scamarcia, Javier Fernandez-Marques, et al. 2025. FlowerTune: A Cross-Domain Benchmark for Federated Fine-Tuning of Large Language Models. https://doi.org/10.48550/ARXIV.2506.02961.
Ghosh, Adhiraj, Sebastian Dziadzio, Ameya Prabhu, Vishaal Udandarao, Samuel Albanie, and Matthias Bethge. 2024. ONEBench to Test Them All: Sample-Level Benchmarking over Open-Ended Capabilities. https://doi.org/10.48550/ARXIV.2412.06745.
Guo, Zishan, Renren Jin, Chuang Liu, et al. 2023. Evaluating Large Language Models: A Comprehensive Survey. https://doi.org/10.48550/ARXIV.2310.19736.
Hamna, Hamna, Gayatri Bhat, Sourabrata Mukherjee, et al. 2026. “Building Benchmarks from the Ground up: Community-Centered Evaluation of LLMs in Healthcare Chatbot Settings.” Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems, April 13, 1–19. https://doi.org/10.1145/3772318.3791172.
Huang, Qian, Chunlan Lv, Li Lu, and Shuang Tu. 2025. “Evaluating the Quality of AI-Generated Digital Educational Resources for University Teaching and Learning.” Systems 13 (3): 174. https://doi.org/10.3390/systems13030174.
Husom, Erik Johannes, Arda Goknil, Merve Astekin, et al. 2025. “Sustainable LLM Inference for Edge AI: Evaluating Quantized LLMs for Energy Efficiency, Output Accuracy, and Inference Latency.” ACM Transactions on Internet of Things 6 (4): 1–35. https://doi.org/10.1145/3767742.
Jegham, Nidhal, Marwan Abdelatti, Chan Young Koh, Lassad Elmoubarki, and Abdeltawab Hendawi. 2025. How Hungry Is AI? Benchmarking Energy, Water, and Carbon Footprint of LLM Inference. https://doi.org/10.48550/ARXIV.2505.09598.
Jeon, YoungHoon, Suwan Kim, Haein Son, Sookbun Lee, Yeil Jeong, and Unggi Lee. 2026. ISD-Agent-Bench: A Comprehensive Benchmark for Evaluating LLM-Based Instructional Design Agents. https://doi.org/10.48550/ARXIV.2602.10620.
Jiang, Yilin, Mingzi Zhang, Xuanyu Yin, et al. 2025. EduGuardBench: A Holistic Benchmark for Evaluating the Pedagogical Fidelity and Adversarial Safety of LLMs as Simulated Teachers. https://doi.org/10.48550/ARXIV.2511.06890.
Jiao, Junfeng, Saleh Afroogh, Abhejay Murali, Kevin Chen, David Atkinson, and Amit Dhurandhar. 2025. “LLM Ethics Benchmark: A Three-Dimensional Assessment System for Evaluating Moral Reasoning in Large Language Models.” Scientific Reports 15 (1). https://doi.org/10.1038/s41598-025-18489-7.
Kim, Eunsu, Haneul Yoo, Guijin Son, Hitesh Patel, Amit Agarwal, and Alice Oh. 2025. BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation. https://doi.org/10.48550/ARXIV.2506.00482.
Lee, Jinsook, Yann Hicke, Renzhe Yu, Christopher Brooks, and René F. Kizilcec. 2024. “The Life Cycle of Large Language Models in Education: A Framework for Understanding Sources of Bias.” British Journal of Educational Technology 55 (5): 1982–2002. https://doi.org/10.1111/bjet.13505.
Lee, Unggi, Hoyoung Ahn, Yoon Choi, et al. 2026. Are Video Models Zero-Shot Learners and Reasoners in Education? EduVideoBench, a Knowledge-Skills-Attitude Benchmark for Educational Video Generation. https://doi.org/10.48550/ARXIV.2605.26918.
Lee, Unggi, Sookbun Lee, Heungsoo Choi, et al. 2026. OpenLearnLM Benchmark: A Unified Framework for Evaluating Knowledge, Skill, and Attitude in Educational Large Language Models. https://doi.org/10.48550/ARXIV.2601.13882.
Lelièvre, Maxime, Amy Waldock, Meng Liu, et al. 2025. Benchmarking the Pedagogical Knowledge of Large Language Models. https://doi.org/10.48550/ARXIV.2506.18710.
Li, Yanhong, Tianyang Xu, Kenan Tang, Karen Livescu, David McAllester, and Jiawei Zhou. 2025. OKBench: Democratizing LLM Evaluation with Fully Automated, on-Demand, Open Knowledge Benchmarking. https://doi.org/10.48550/ARXIV.2511.08598.
Lin, Pingrong, Qin Deng, and Yanbian Zhou. 2026. “Towards Responsible AI in Education: A Delphi-AHP-Based Framework for Evaluating Educational Large Language Models.” Computers and Education: Artificial Intelligence 10 (June): 100534. https://doi.org/10.1016/j.caeai.2025.100534.
Liu, Hongwei, Zilong Zheng, Yuxuan Qiao, et al. 2024. MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark. https://doi.org/10.48550/ARXIV.2405.12209.
Lunardi, Riccardo, Vincenzo Della Mea, Stefano Mizzaro, and Kevin Roitero. 2025. On Robustness and Reliability of Benchmark-Based Evaluation of LLMs. https://doi.org/10.48550/ARXIV.2509.04013.
Martins, Pedro Henrique, Patrick Fernandes, João Alves, et al. 2024. EuroLLM: Multilingual Language Models for Europe. https://doi.org/10.48550/ARXIV.2409.16235.
McIntosh, Timothy R., Teo Susnjak, Nalin Arachchilage, et al. 2026. “Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence.” IEEE Transactions on Artificial Intelligence 7 (1): 22–39. https://doi.org/10.1109/tai.2025.3569516.
Mehditabar, Mohammadjavad, Saurabhsingh Rajput, Antonio Mastropaolo, and Tushar Sharma. 2025. Smart but Costly? Benchmarking LLMs on Functional Accuracy and Energy Efficiency. https://doi.org/10.48550/ARXIV.2511.07698.
Moëll, Birger, Fabian Farestam, and Jonas Beskow. 2025. “Swedish Medical LLM Benchmark: Development and Evaluation of a Framework for Assessing Large Language Models in the Swedish Medical Domain.” Frontiers in Artificial Intelligence 8 (July). https://doi.org/10.3389/frai.2025.1557920.
Mohammadi, Mahmoud, Yipeng Li, Jane Lo, and Wendy Yip. 2025. “Evaluation and Benchmarking of LLM Agents: A Survey.” Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2, August 3, 6129–39. https://doi.org/10.1145/3711896.3736570.
Nikiforova-Ilieva, Kalina, and Tsvetozar Georgiev. 2025. “Benchmarking and Evaluation Framework for Large Language Models in Education.” 2025 6th International Conference on Communications, Information, Electronic and Energy Systems (CIEES), November 26, 1–5. https://doi.org/10.1109/ciees66347.2025.11300167.
Noels, Sander, Jorne De Blaere, and Tijl De Bie. 2024. “A Dutch Financial Large Language Model.” Proceedings of the 5th ACM International Conference on AI in Finance, November 14, 283–91. https://doi.org/10.1145/3677052.3698628.
Polo, Felipe Maia, Lucas Weber, Leshem Choshen, Yuekai Sun, Gongjun Xu, and Mikhail Yurochkin. 2024. tinyBenchmarks: Evaluating LLMs with Fewer Examples. https://doi.org/10.48550/ARXIV.2402.14992.
Pomerenke, David, Jonas Nothnagel, and Simon Ostermann. 2025. The AI Language Proficiency Monitor – Tracking the Progress of LLMs on Multilingual Benchmarks. https://doi.org/10.48550/ARXIV.2507.08538.
Potamitis, Nearchos, Vansh Ramani, Har Ashish Arora, Dhairya Kuchhal, Lars Klein, and Akhil Arora. 2025. ReasonBENCH: Benchmarking the (in)stability of LLM Reasoning. https://doi.org/10.48550/ARXIV.2512.07795.
Qian, Qi, Chengsong Huang, Jingwen Xu, et al. 2026. Benchmark^2: Systematic Evaluation of LLM Benchmarks. https://doi.org/10.48550/ARXIV.2601.03986.
Quttainah, Majdi, Vinaytosh Mishra, Somayya Madakam, Yotam Lurie, and Shlomo Mark. 2024. “Cost, Usability, Credibility, Fairness, Accountability, Transparency, and Explainability Framework for Safe and Effective Large Language Models in Medical Education: Narrative Review and Qualitative Study.” JMIR AI 3 (April): e51834. https://doi.org/10.2196/51834.
Rooein, Donya, Paul Rottger, Anastassia Shaitarova, and Dirk Hovy. 2024. Beyond Flesch-Kincaid: Prompt-Based Metrics Improve Difficulty Classification of Educational Texts. https://doi.org/10.48550/ARXIV.2405.09482.
Shahzad, Tariq, Tehseen Mazhar, Muhammad Usman Tariq, Wasim Ahmad, Khmaies Ouahada, and Habib Hamam. 2025. “A Comprehensive Review of Large Language Models: Issues and Solutions in Learning Environments.” Discover Sustainability 6 (1). https://doi.org/10.1007/s43621-025-00815-8.
Shashidhar, Sumuk, Clémentine Fourrier, Alina Lozovskia, Thomas Wolf, Gokhan Tur, and Dilek Hakkani-Tür. 2025. YourBench: Easy Custom Evaluation Sets for Everyone. https://doi.org/10.48550/ARXIV.2504.01833.
Shool, Sina, Sara Adimi, Reza Saboori Amleshi, Ehsan Bitaraf, Reza Golpira, and Mahmood Tara. 2025. “A Systematic Review of Large Language Model (LLM) Evaluations in Clinical Medicine.” BMC Medical Informatics and Decision Making 25 (1). https://doi.org/10.1186/s12911-025-02954-4.
Singh, Shivalika, Yiyang Nan, Alex Wang, et al. 2025. The Leaderboard Illusion. https://doi.org/10.48550/ARXIV.2504.20879.
Son, Guijin, Dongkeun Yoon, Juyoung Suk, et al. 2024. MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models. https://doi.org/10.48550/ARXIV.2410.17578.
Thellmann, Klaudia, Bernhard Stadler, Michael Fromm, et al. 2024. Towards Multilingual LLM Evaluation for European Languages. https://doi.org/10.48550/ARXIV.2410.08928.
Vanroy, Bram. 2023. Language Resources for Dutch Large Language Modelling. https://doi.org/10.48550/ARXIV.2312.12852.
Vanroy, Bram. 2024. Fietje: An Open, Efficient LLM for Dutch. https://doi.org/10.48550/ARXIV.2412.15450.
Vijay, Dhaathri, and Anandaswarup Vadapalli. 2025. The Hidden Costs of Translation Accuracy: Distillation, Quantization, and Environmental Impact. https://doi.org/10.48550/ARXIV.2509.23990.
Vintar, Špela, Taja Kuzman Pungeršek, Mojca Brglez, and Nikola Ljubešić. 2025. Charting the European LLM Benchmarking Landscape: A New Taxonomy and a Set of Best Practices. https://doi.org/10.48550/ARXIV.2510.24450.
Vries, Wietse de, Martijn Wieling, and Malvina Nissim. 2023. DUMB: A Benchmark for Smart Evaluation of Dutch Models. https://doi.org/10.48550/ARXIV.2305.13026.
Wang, Shaobo, Cong Wang, Wenjie Fu, et al. 2025. Rethinking LLM Evaluation: Can We Evaluate LLMs with 200x Less Data? https://doi.org/10.48550/ARXIV.2510.10457.
Wei, Shou’ang, Xinyun Wang, Shuzhen Bi, et al. 2025. ELMES: An Automated Framework for Evaluating Large Language Models in Educational Scenarios. https://doi.org/10.48550/ARXIV.2507.22947.
Wu, Yanran, Inez Hua, and Yi Ding. 2025. Unveiling Environmental Impacts of Large Language Model Serving: A Functional Unit View. https://doi.org/10.48550/ARXIV.2502.11256.
Xu, Cheng, Shuhao Guan, Derek Greene, and M-Tahar Kechadi. 2024. Benchmark Data Contamination of Large Language Models: A Survey. https://doi.org/10.48550/ARXIV.2406.04244.
Yuan, Zhengqing, Weixiang Sun, Yixin Liu, et al. 2025. EfficientLLM: Efficiency in Large Language Models. https://doi.org/10.48550/ARXIV.2505.13840.
Zhang, Dan, Sining Zhoubian, Min Cai, et al. 2025. DataSciBench: An LLM Agent Benchmark for Data Science. https://doi.org/10.48550/ARXIV.2502.13897.
Zhou, Chengliang, Mei Wang, Ting Zhang, Qiannan Zhu, Jian Li, and Hua Huang. 2025. From Answers to Questions: EQGBench for Evaluating LLMs’ Educational Question Generation. https://doi.org/10.48550/ARXIV.2508.10005.
Zhu, Longyuan, Hairan Hua, Linlin Miao, and Bing Zhao. 2026. Benchmark Health Index: A Systematic Framework for Benchmarking the Benchmarks of LLMs. https://doi.org/10.48550/ARXIV.2602.11674.