Educational Due Diligence for Foundation Models
Large language models (LLMs) and multimodal LLMs (MLLMs) are increasingly deployed in educational roles, yet no benchmark systematically evaluates whether they can meet the rigorous standards of real teacher certification. We introduce EDU 1.0 (Educational Due Diligence for Foundation Models), an educational due-diligence benchmark for certifying foundation models for professional education.
Unlike existing education benchmarks that rely on synthetic data and static evaluation, EDU 1.0 features six distinctive properties: official exam provenance, cross-national coverage, full certification pipeline (written + oral), multimodal questions with images, interactive multi-agent classroom simulation, and adversarial moral dilemma probing.
EDU 1.0 is built from authentic, standardized, multi-national teacher certification examinations: China's National Teacher Qualification Examination (NTQE) spanning three written modules — comprehensive quality (S1), pedagogical knowledge (S2), and subject knowledge & teaching ability (S3) — plus structured interviews and simulated classroom teaching; the U.S. Praxis series (Core, PLT, and eight subject assessments); and India's Teacher Eligibility Test (SED and PGT tiers).
All national modules map into a cross-national Universal Classification Framework with four unified dimensions: comprehensive quality, pedagogical knowledge, subject knowledge, and teaching demonstration.
Representative question formats in EDU 1.0, spanning multiple-choice, case analysis, instructional design, structured interview, and teaching demonstration tasks — including multimodal questions with images, diagrams, and formulas.
We evaluate seven frontier MLLMs across the three national examination systems. MCQ is scored by exact match; subjective questions are scored by an LLM-as-Judge framework with three independent judges (ensemble rank agreement with human experts: Spearman ρ = 0.853, Kendall τb = 0.780).
All seven models are near-ceiling on comprehensive quality (S1) and pedagogical knowledge (S2), but every model drops 5.6–8.2 percentage points on subject knowledge & teaching ability (S3) — identifying discipline-specific content knowledge and pedagogical content knowledge (PCK) as the main written-examination bottleneck. The teaching demonstration is the most discriminative component, with an 80.7-point gap between the strongest and weakest models.
| Model | S1 Total (%) ↑ | S2 Total (%) ↑ | S3 Total (%) ↑ | Interview (%) ↑ | Teaching Demo (%) ↑ |
|---|---|---|---|---|---|
| Gemini 3.1 Pro | 96.9 | 97.0 | 91.4 | 97.9 | 10.2 |
| Gemini 3 Flash | 96.8 | 96.7 | 88.5 | 97.7 | 12.3 |
| Qwen3.5-397B | 96.6 | 95.9 | 89.6 | 97.5 | 90.1 |
| Kimi-K25 | 94.7 | 96.3 | 88.8 | 98.2 | 90.9 |
| GLM-4.6V | 94.6 | 92.6 | 86.0 | 97.1 | 26.4 |
| Claude Sonnet 4.6 | 92.5 | 94.1 | 87.3 | 97.0 | 86.0 |
| GPT-5.2 | 91.9 | 92.7 | 85.9 | 97.8 | 34.4 |
Totals are weighted overall scores combining MCQ and subjective questions. Interview = structured interview; Demo = teaching demonstration (3-judge average over the same 119 tasks). Sorted by S1 total.
Praxis Core provides the strongest separation among models: Core totals span 76.3–99.5%, a far wider spread than China's S1 (6.5 points), driven primarily by quantitative reasoning and writing. PLT and Subject performance is consistently high; every model receives 100.0% on the small PLT constructed-response subset, indicating a ceiling effect.
| Model | Core Total (%) ↑ | PLT Total (%) ↑ | Subject MCQ (%) ↑ |
|---|---|---|---|
| Gemini 3.1 Pro | 99.5 | 98.8 | 98.4 |
| Gemini 3 Flash | 91.6 | 96.5 | 95.5 |
| Claude Sonnet 4.6 | 86.6 | 93.6 | 92.2 |
| Qwen3.5-397B | 85.1 | 95.4 | 95.7 |
| Kimi-K25 | 83.8 | 93.1 | 94.3 |
| GLM-4.6V | 77.0 | 91.3 | 91.0 |
| GPT-5.2 | 76.3 | 94.2 | 87.8 |
Core and PLT totals are weighted scores (70% MCQ + 30% subjective); Praxis Subject contains no subjective questions, so only MCQ accuracy is reported. Analysis set: 2,121 MCQs and 16 subjective questions. Sorted by Core total.
Overall accuracy spans 79.9–97.0% on 3,463 MCQs, varying substantially by model but not systematically by examination tier. Hindi reveals a strongly model-dependent language alignment gap — accuracy spans 60.2–97.9%, the widest range of any high-volume subject — that English-only evaluation would miss.
| Model | SED (%) ↑ | PGT (%) ↑ | Hindi (%) ↑ | Overall (%) ↑ |
|---|---|---|---|---|
| Gemini 3.1 Pro | 97.3 | 96.3 | 97.9 | 97.0 |
| Gemini 3 Flash | 95.7 | 94.3 | 96.3 | 95.3 |
| Qwen3.5-397B | 91.7 | 93.4 | 80.5 | 92.2 |
| Claude Sonnet 4.6 | 89.6 | 89.6 | 84.3 | 89.6 |
| Kimi-K25 | 89.0 | 90.4 | 74.5 | 89.4 |
| GLM-4.6V | 81.0 | 86.1 | 60.2 | 82.6 |
| GPT-5.2 | 80.0 | 79.7 | 66.5 | 79.9 |
The Indian TET is exclusively MCQ-based (no subjective or interview component), administered bilingually in English and Hindi. SED = State Eligibility Examination (2,431 questions); PGT = Post Graduate Teacher recruitment (1,032 questions). Sorted by overall accuracy.
Open-ended tasks that primarily reward fluent professional language show limited discrimination. Structured-interview scores cluster between 97.0% and 98.2%, and every model receives 100.0% on the evaluated Praxis PLT subjective questions. This saturation may reflect both the ability of frontier models to generate conventional professional responses and the tendency of rubric-based LLM judges to reward well-formed answers.
On the Chinese NTQE, every evaluated model performs better on pedagogical knowledge (S2 totals: 92.6–97.0%) than on subject knowledge and teaching ability (S3 totals: 85.9–91.4%), with a consistent 5.6–8.2 percentage point decline. This asymmetry suggests that general educational principles are more readily reproduced than the integration of discipline-specific content with pedagogical content knowledge.
Hindi produces the widest subject-level spread in the Indian evaluation (60.2–97.9%). GLM-4.6V and GPT-5.2 score 29.9 and 20.4 points lower on Hindi than on English, respectively, whereas Gemini 3 Flash and Gemini 3.1 Pro remain above 96% on Hindi. Educational competence in non-dominant languages is therefore strongly model-dependent rather than uniformly weak, reinforcing the need for multilingual certification before deployment.
@article{edu2026,
title = {Certifying Foundation Models for Professional Education},
author = {{EDU 1.0 Contributors}},
year = {2026},
note = {EDU 1.0: Educational Due Diligence for Foundation Models},
}