EDU 1.0

Educational Due Diligence for Foundation Models

📄 Paper 🗂 Dataset 🏆 Leaderboard GitHub

Latest News

Introduction

Large language models (LLMs) and multimodal LLMs (MLLMs) are increasingly deployed in educational roles, yet no benchmark systematically evaluates whether they can meet the rigorous standards of real teacher certification. We introduce EDU 1.0 (Educational Due Diligence for Foundation Models), an educational due-diligence benchmark for certifying foundation models for professional education.

Unlike existing education benchmarks that rely on synthetic data and static evaluation, EDU 1.0 features six distinctive properties: official exam provenance, cross-national coverage, full certification pipeline (written + oral), multimodal questions with images, interactive multi-agent classroom simulation, and adversarial moral dilemma probing.

Dataset

EDU 1.0 is built from authentic, standardized, multi-national teacher certification examinations: China's National Teacher Qualification Examination (NTQE) spanning three written modules — comprehensive quality (S1), pedagogical knowledge (S2), and subject knowledge & teaching ability (S3) — plus structured interviews and simulated classroom teaching; the U.S. Praxis series (Core, PLT, and eight subject assessments); and India's Teacher Eligibility Test (SED and PGT tiers).

10,593
questions (NTQE + Praxis + TET)
13
secondary subjects
5
certification stages

All national modules map into a cross-national Universal Classification Framework with four unified dimensions: comprehensive quality, pedagogical knowledge, subject knowledge, and teaching demonstration.

Examples

China · NTQE S3 — Subject Knowledge & Teaching Ability (Mathematics)
Instructional design: Based on the lesson "Properties of Quadratic Functions" for Grade 9, design a 10-minute teaching segment that includes clear learning objectives, a guided-discovery activity, and at least one formative assessment checkpoint. Explain how your design addresses common student misconceptions about the vertex form.
Scored by a three-judge LLM-as-Judge ensemble against official rubrics.
United States · Praxis PLT (Grades 7–12)
Question: A teacher notices that students consistently perform well on recall items but struggle with questions requiring transfer to novel contexts. Which instructional adjustment is most consistent with constructivist learning theory? (Four-option MCQ drawn from official ETS study companions.)
Exact-match scoring against the standard answer key.

Representative question formats in EDU 1.0, spanning multiple-choice, case analysis, instructional design, structured interview, and teaching demonstration tasks — including multimodal questions with images, diagrams, and formulas.

Quantitative Results

We evaluate seven frontier MLLMs across the three national examination systems. MCQ is scored by exact match; subjective questions are scored by an LLM-as-Judge framework with three independent judges (ensemble rank agreement with human experts: Spearman ρ = 0.853, Kendall τb = 0.780).

China NTQE — Written Modules & Oral Stage

All seven models are near-ceiling on comprehensive quality (S1) and pedagogical knowledge (S2), but every model drops 5.6–8.2 percentage points on subject knowledge & teaching ability (S3) — identifying discipline-specific content knowledge and pedagogical content knowledge (PCK) as the main written-examination bottleneck. The teaching demonstration is the most discriminative component, with an 80.7-point gap between the strongest and weakest models.

Model S1 Total (%) ↑ S2 Total (%) ↑ S3 Total (%) ↑ Interview (%) ↑ Teaching Demo (%) ↑
Gemini 3.1 Pro96.997.091.497.910.2
Gemini 3 Flash96.896.788.597.712.3
Qwen3.5-397B96.695.989.697.590.1
Kimi-K2594.796.388.898.290.9
GLM-4.6V94.692.686.097.126.4
Claude Sonnet 4.692.594.187.397.086.0
GPT-5.291.992.785.997.834.4

Totals are weighted overall scores combining MCQ and subjective questions. Interview = structured interview; Demo = teaching demonstration (3-judge average over the same 119 tasks). Sorted by S1 total.

United States Praxis — Core, PLT & Subject Assessments

Praxis Core provides the strongest separation among models: Core totals span 76.3–99.5%, a far wider spread than China's S1 (6.5 points), driven primarily by quantitative reasoning and writing. PLT and Subject performance is consistently high; every model receives 100.0% on the small PLT constructed-response subset, indicating a ceiling effect.

Model Core Total (%) ↑ PLT Total (%) ↑ Subject MCQ (%) ↑
Gemini 3.1 Pro99.598.898.4
Gemini 3 Flash91.696.595.5
Claude Sonnet 4.686.693.692.2
Qwen3.5-397B85.195.495.7
Kimi-K2583.893.194.3
GLM-4.6V77.091.391.0
GPT-5.276.394.287.8

Core and PLT totals are weighted scores (70% MCQ + 30% subjective); Praxis Subject contains no subjective questions, so only MCQ accuracy is reported. Analysis set: 2,121 MCQs and 16 subjective questions. Sorted by Core total.

India TET — State Eligibility (SED) & PGT Recruitment

Overall accuracy spans 79.9–97.0% on 3,463 MCQs, varying substantially by model but not systematically by examination tier. Hindi reveals a strongly model-dependent language alignment gap — accuracy spans 60.2–97.9%, the widest range of any high-volume subject — that English-only evaluation would miss.

Model SED (%) ↑ PGT (%) ↑ Hindi (%) ↑ Overall (%) ↑
Gemini 3.1 Pro97.396.397.997.0
Gemini 3 Flash95.794.396.395.3
Qwen3.5-397B91.793.480.592.2
Claude Sonnet 4.689.689.684.389.6
Kimi-K2589.090.474.589.4
GLM-4.6V81.086.160.282.6
GPT-5.280.079.766.579.9

The Indian TET is exclusively MCQ-based (no subjective or interview component), administered bilingually in English and Hindi. SED = State Eligibility Examination (2,431 questions); PGT = Post Graduate Teacher recruitment (1,032 questions). Sorted by overall accuracy.

Key Findings

Subjective ceiling effects in open-ended language tasks

Open-ended tasks that primarily reward fluent professional language show limited discrimination. Structured-interview scores cluster between 97.0% and 98.2%, and every model receives 100.0% on the evaluated Praxis PLT subjective questions. This saturation may reflect both the ability of frontier models to generate conventional professional responses and the tendency of rubric-based LLM judges to reward well-formed answers.

The pedagogy–subject knowledge gap

On the Chinese NTQE, every evaluated model performs better on pedagogical knowledge (S2 totals: 92.6–97.0%) than on subject knowledge and teaching ability (S3 totals: 85.9–91.4%), with a consistent 5.6–8.2 percentage point decline. This asymmetry suggests that general educational principles are more readily reproduced than the integration of discipline-specific content with pedagogical content knowledge.

Model-dependent language alignment

Hindi produces the widest subject-level spread in the Indian evaluation (60.2–97.9%). GLM-4.6V and GPT-5.2 score 29.9 and 20.4 points lower on Hindi than on English, respectively, whereas Gemini 3 Flash and Gemini 3.1 Pro remain above 96% on Hindi. Educational competence in non-dominant languages is therefore strongly model-dependent rather than uniformly weak, reinforcing the need for multilingual certification before deployment.

Citation

@article{edu2026,
  title  = {Certifying Foundation Models for Professional Education},
  author = {{EDU 1.0 Contributors}},
  year   = {2026},
  note   = {EDU 1.0: Educational Due Diligence for Foundation Models},
}