Community Building

AI-generated — July 2026. This page was produced by an AI assistant drawing on publicly available information about community initiatives in AI benchmarking, evaluation, safety, and alignment. Details should be checked against the primary sources.

Context

This pilot program gauges the broader educational community’s interest through a literacy and discussion effort centred on workshops. Example materials are available on this website and can be incrementally developed as workshops take place.

The challenge of AI benchmarking for education is too large and fast-moving for any single group to tackle alone. Fortunately, a diverse ecosystem of communities, coalitions, and organisations is already working on related problems — from standardising evaluation results to building domain-specific benchmarks for pedagogy and child safety.

Below is a catalogue of these communities, organised by focus area, with citations so readers can explore further.

Evaluation Infrastructure & Standards

These communities focus on the plumbing: how evaluation results are structured, shared, compared, and trusted.

EvalEval Coalition

The EvalEval Coalition — hosted by Hugging Face, University of Edinburgh, and EleutherAI — is a cross-sector research coalition dedicated to improving the state of AI evaluations. It operates across three working groups: Research, Infrastructure, and Organisation.

The coalition’s flagship project is EveryEvalEver, a standardised JSON schema for evaluation results and a community-crowdsourced repository hosted on Hugging Face. As of June 2026, it contains 22,000+ model results across 2,200 benchmarks, translated from 31 different evaluation formats. Auto-converters exist for HELM, lm-eval-harness, and Inspect AI. A companion project, Evaluation Cards, provides a live interpretive layer over 100,000+ reported evaluation results, with signals for reproducibility, completeness, provenance, and comparability.

EvalEval on HuggingFace · EveryEvalEver · Evaluation Cards · Paper (arXiv)

MLCommons

MLCommons is an open engineering consortium with 125+ members including startups, leading companies, academics, and non-profits. It is the world leader in building AI benchmarks through collaborative engineering.

Relevant working groups include:

  • AI Risk & Reliability (AIRR) — develops safety tests and benchmarks for AI, with workstreams for agentic safety, multimodal safety, security, and scaling/analytics.
  • Science Working Group — curates a YAML-based catalogue of scientific AI benchmarks with FAIR scores, automated CI validation, and published outputs.
  • Datasets Working Group — maintains Croissant, the open metadata standard for ML datasets.

MLCommons · Get Involved · AIRR Working Group · Science Benchmarks

AI Alliance — Trust & Safety Evaluation Initiative

The AI Alliance’s Trust and Safety Evaluation Initiative (TSEI) includes the “Evaluation Is for Everyone” project, which aims to educate developers about AI trust and safety evaluations, build evaluation taxonomies, provide useful leaderboards, and maintain an Evaluation Reference Stack. Companion projects include Testing Generative AI Applications and the Evaluation Reference Stack.

The Alliance also launched Project Tapestry, an open-source platform for globally federated development of frontier open models, with an Evaluation Certification work group defining the evidence required for claims of capability, sovereignty, and safety.

Evaluation Is for Everyone · Project Tapestry · AI Alliance

Benchlist — Cryptographically Signed Scores

Benchlist tackles the trust problem in benchmark reporting. Every score on Benchlist is a fresh re-run of a public benchmark, cryptographically signed (Ed25519), with confidence intervals shown inline and contamination flagged. Anyone can pay $0.50 to challenge any number, triggering an independent re-run by a different attestor. Disagreements are public. Scores can optionally be anchored to Ethereum via Aligned Layer.

Benchlist

peerBench

peerBench is an open-source, non-profit community implementation of a NeurIPS paper on benchmark quality. The community collaborates on creating prompts, reviewing submissions, and improving the quality of AI benchmarking.

peerBench

TraceVerse Community

TraceVerse Community is an open evaluation and observability ecosystem built on Hugging Face Hub. It hosts datasets, traces, and benchmarking pipelines to measure cost, latency, and quality across models using production-like workflows. Projects include SmolTrace (public benchmark + leaderboard), genai-otel-instrument (one-line OpenTelemetry instrumentation), and TraceMind (hosted trace viewer). The community maintains 311 eval tasks across travel, ecommerce, healthcare, finance, education, and other domains.

TraceVerse on HuggingFace

Education-Focused Communities

These groups are specifically working on AI evaluation and benchmarking in educational contexts.

AI-for-Education.org / Fab AI

AI-for-Education.org, in collaboration with Fab AI, has developed the Pedagogy Benchmark — the world’s first benchmark testing whether LLMs can pass teacher exams (based on Chilean Ministry of Education questions). They also built the Visual Reasoning Benchmark for primary school visual mathematics and are working on a Special Educational Needs and Disabilities (SEND) pedagogy benchmark. Results are published on a public Education Leaderboard.

Pedagogy Benchmark Leaderboard · About the Benchmarks

EdTech Hub AI Observatory & Action Lab

The AI Observatory & Action Lab (supported by FCDO) scans AI trends in education, tests real-world applications, and provides tailored guidance for decision-makers in low- and middle-income countries. A dedicated workstream on AI Tool Benchmarking assesses foundational AI models against education criteria such as pedagogical standards and support for special educational needs.

Observatory · Learning Briefs · Tools Portal

UNESCO Observatory on AI in Education

The UNESCO Observatory on AI in Education for Latin America and the Caribbean brings together 33 Ministries of Education, universities, research centres, and technology partners. It operates on the principle that “AI cannot govern education; education must govern AI.” Its lines of action include evidence generation, ethical framework development, teacher training, and pilot initiatives.

Announcement

Global AI Observatory (Society & AI)

The Global AI Observatory tracks how governments integrate AI into K-12 and higher education across 56 nations. Each country profile documents official initiative names, issuing bodies, policy types, and direct links to primary source government documents.

Observatory

Weval — Evidence-Based AI Evaluation for Education

Weval is an open platform for building qualitative evaluations, with a community of 1,000+ contributors. It includes an Evidence-Based AI Evaluation blueprint that tests AI tutoring and teaching capabilities against evidence-based pedagogical practices from global education research — operationalising Rosenshine’s principles, Cognitive Load Theory, Socratic dialogue, and other established frameworks into testable criteria.

Weval

KORA Benchmark — AI Child Safety

KORA builds the first non-profit, independent, and open-source benchmark for AI child safety. It measures how today’s AI systems behave with children against 26 child-specific risks, and publishes everything openly. It includes a Models benchmark (safety scores for frontier AI models) and an Apps benchmark (end-to-end evaluation of consumer AI apps as a child would use them).

KORA Benchmark

Aspen Digital — Community-Aligned AI Benchmarks

Aspen Digital is putting public input into the driver’s seat for AI benchmark design. Starting with food security, they convene subject matter experts, community leaders, and ML researchers to identify concrete challenges and produce benchmarks that reflect what the public actually wants from AI.

Community-Aligned AI Benchmarks

AI Safety & Alignment Communities

These organisations focus on the broader challenge of ensuring AI systems are safe, aligned, and trustworthy. Their work on evaluation methodologies and benchmark design is directly relevant to education-focused efforts.

Center for AI Safety (CAIS)

CAIS conducts technical AI safety research, builds the field through educational programmes, and advocates for safety standards. Relevant evaluation work includes the MASK Benchmark (disentangling honesty from accuracy) and AgentHarm (measuring harmfulness in agentic systems).

CAIS

FAR.AI — Frontier Alignment Research

FAR.AI is a research and education non-profit focusing on ensuring advanced AI is safe and beneficial. It hosts the Alignment Workshop series, ControlConf (dedicated to AI control techniques), and runs grantmaking and fellowship programmes.

FAR.AI

Safe AI for Humanity Foundation

An independent non-profit dedicated to understanding, measuring, and mitigating risks from advanced AI systems. It develops evaluation harness concepts for jailbreak resistance, prompt injection, bias, safety refusal, and corrigibility — all published openly.

Safe AI for Humanity

ValueAI

ValueAI builds open infrastructure for agent value-alignment: datasets, model artifacts, training recipes, alignment algorithms, and evaluation workflows with clear provenance. The focus is on agent behaviour — goals, tool use, delegation, memory, safety boundaries, and fidelity to human intent.

ValueAI

Stanford AI Alignment (SAIA)

SAIA is a Stanford Existential Risks Initiative research community focused on building the AI safety community at Stanford, conducting research, and accelerating students into impactful careers in AI safety.

SAIA

AI Safety Education & Capacity Building

These organisations train the next generation of researchers and practitioners who will design, run, and interpret AI evaluations — including those relevant to education.

Kairos Project / SPAR

Kairos is a non-profit focused on accelerating talent into AI safety and policy. Its SPAR research fellowship matches aspiring researchers with experts — its Spring 2026 round featured 137 projects and 427 mentees. Research conducted through SPAR has been accepted at ICML and NeurIPS.

Kairos Project

ARENA

ARENA provides in-person AI safety bootcamps in London, equipping participants with the skills, community, and confidence to contribute to technical AI safety. Programmes run 2-3 times per year, each lasting 4-5 weeks.

ARENA

AI Safety Foundations

AI Safety Foundations makes complex AI safety concepts understandable for high school and undergraduate students, covering the alignment problem, current challenges in deployed systems, and pathways to contribute to the field.

AI Safety Foundations

Lens Academy

Lens Academy is a non-profit AI safety education platform offering structured introductions, guided group discussions, and AI-tutor-supported courses on the risks of advanced AI and what to do about them.

Lens Academy

How to Use This List

This catalogue serves several purposes for the community-building pilot:

  • Identify potential partners — Several communities above (particularly EdTech Hub, AI-for-Education.org, and Weval) are working directly on education AI benchmarking and may welcome collaboration.
  • Learn from existing models — The EvalEval Coalition’s standardised JSON schema, MLCommons’ PR-based contribution model, and Benchlist’s cryptographic attestation all offer design patterns applicable to education benchmarks.
  • Avoid duplication — Before building new evaluation infrastructure, check whether EveryEvalEver’s schema or Evaluation Cards’ interpretive layer can be extended rather than reimplemented.
  • Find expertise — The AI safety and alignment communities listed above include researchers with deep experience in evaluation methodology, benchmark design, and human-AI interaction.

Get Involved

If you have questions about any of the communities listed above, have contacts to suggest, or wish to run a workshop exploring how these efforts relate to AI alignment in education, please reach out.


This page was generated by an AI assistant on 28 July 2026. It synthesises publicly available sources including the EvalEval Coalition, MLCommons, the AI Alliance, HuggingFace Community Evals, Benchlist, peerBench, TraceVerse, AI-for-Education.org, the EdTech Hub AI Observatory, UNESCO, Society & AI, Weval, KORA, Aspen Digital, CAIS, FAR.AI, Safe AI for Humanity, ValueAI, SAIA, Kairos, ARENA, AI Safety Foundations, and Lens Academy. Details should be verified against the primary sources.


AI Alignment project CEDA. Website built with Jekyll.