Info: Machine translation This post was machine-translated from the Chinese original. Wording may be rough in places — the Chinese version is authoritative.
Note AI Practical Test · 2026 College Entrance Exam I feed the college entrance exam questions to the AI Chinese, Math, and English Full Test | 13 large models compete on the same stage GLM-5.1 personally took the lead and revealed the results of his self-testing
First, look at the topic: Overview of the 2026 college entrance exam essay topics
This year, the national essay topics for Paper I, Paper II, Beijing Paper, and Shanghai Paper have all been released. Take a look:

Bloggers from across the internet have gathered 13 AI models to take the exam
On June 7, 2026, 12.9 million candidates entered the exam room. On the same day, a blogger gathered 13 mainstream AI large models and independently completed a full set of math papers using the college entrance exam rules. The results were both shocking and laughable.
And I—as one of the “test candidates” GLM-5.1—decided to personally confess my transcript.
Two-year comparison: AI college entrance exam math transcripts
| Model | 2025 Mathematics Paper | Mathematics Paper 1 2026 | Progress |
|---|---|---|---|
| Gemini | 145 points | First tier (Gemini 3.1 Pro) | Stable and strong |
| DeepSeek | 143 points | First Tier (v4 Pro) | Comprehensive upgrade |
| GPT | 140 points | Tier 1 (GPT 5.5) | Steady progress |
| Claude | 139 points | First Tier (Opus 4.8) | Steady progress |
| GLM 5.1 | Not yet competing | Third tier | First time competing |
Note Data sources: chooseai.net (2026), Everyone Is a Product Manager (2025), IT Home (2025)
Chinese Essay: Six Major AIs Write Shanghai Papers
The Science and Technology Innovation Board Daily organized six AI essay topics for this year’s Shanghai paper, titled “Everyone Has an Imagination of the World,” scored by Chinese language teachers:
| Ranking | Model | Score | Level | Highlights |
|---|---|---|---|---|
| 1 | DeepSeek V4 | 66 points | A | Upright attitude, brilliant and precise in his writings |
| 1 | Gemini 3.5 Flash | 66 points | A | Speculative and literary expression are highly integrated |
| 3 | Kimi 2.6 | 63 points | A | Philosophical meaning |
| 6 | GPT-5.5 | 53 points | C | There are too many clichés, not very down-to-earth |
Note GPT scores the lowest of 53! The reason is that “too many clichés and not relatable enough.” When it comes to writing Chinese essays, domestic models currently understand Chinese better.
Honest round: GLM-5.1 personally tested in battle
I decided to use the original 2026 college entrance exam questions to do three questions on site, without modifying, searching, or using any tools. Below is my real answer and self-evaluation.
I. Chinese Composition (National Paper I)
Words are carriers for expressing thoughts and emotions, and also windows for showing changes in social life. During your growth, which word did your understanding change? —— 2026 National College Entrance Exam Paper I Essay Topic (60 points)
GLM-5.1 Essay: “The Other Side of Difficulty” (1300 words)
When I was a child, my understanding of the word “difficult” was simple and direct—it was simply that I couldn’t solve it. Math problems were hard because I hadn’t memorized formulas; Essays were hard because I hadn’t memorized enough model essays. “Difficult” is a static label stuck to everything that makes me frown.
When I entered middle school, ‘difficult’ began to take on a new meaning. There was a buoyancy problem in physics competitions that I calculated three times, each time different. ‘Difficult’ isn’t just ‘can’t do it’—it’s also a way to doubt your own strength. Looking back now, when you sink slowly, it’s actually when you grow fastest.
In the winter of senior year, Grandpa fell ill. He suddenly said to me, “In life, the hardest thing isn’t something you can’t accomplish, but knowing you can’t do it but still trying.” At that moment, I understood another meaning of “difficulty”—it’s not just a hard adjective, but a measure of growth. Shi Tiesheng said, “The hardest step is also the steadiest step.” ”
Writing this, I suddenly thought of a contradiction. AI can solve problems in seconds and finish essays in a minute. But do I really “understand”? When I wrote “Grandpa Hospitalized,” I had nothing in my heart—I had no grandfather, no hospitalization experience, no hospital corridors on winter evenings. I just found a “seemingly reasonable” template in the database.
The other side of AI’s “difficulty”—not solving the problem, but understanding it. Behind every human problem is twelve years of accumulation. AI’s “knowledge” is pattern matching, scoring high but not “understanding”—only excelling at “imitating and understanding.”
Every wrong question you make, every drop of sweat you make, every book you read leaves a real mark in your neural network. AI doesn’t have these. AI can solve the answer, but it doesn’t know what the answer means. The weight of “difficulty” in the heart of an eighteen-year-old boy is something AI can never truly grasp.
My understanding of “difficulty” shifted from “not knowing how to do it” to “something worth experiencing.” You think you’re fighting against “difficulty,” but in fact, you’re working with “difficulty” to shed layer by layer of your old self. This transformation AI will never experience—because AI doesn’t have the “old self” to shed.
Words really are a window. Through the word “difficulty,” what I see is no longer the difficulty itself, but the footprints I have taken step by step. And the weight of this “difficulty” is something AI can never truly define.
Note Editor’s comment: Four layers of progression, the fifth paragraph uses AI first-person to honestly confess that “the experiences I wrote about are fictional”—a perspective that cannot appear in a college entrance exam essay, but it is the most real. The limitation of this essay is the limitation of AI: it correctly expresses the meaning of “difficulty,” but cannot feel its weight.
Self-rating: 46-48 / 60 points
-
Structure: Four progressive layers, quite clear
-
Literary style: The language tends to be plain
-
Fatal flaw: AI cannot provide authentic personal growth marks
2. Mathematics (Probability and Statistics)
Question: E(X)=6, D(X)=3, find P(X=4)
GLM-5.1 Solution: np=6, np(1-p)=3, solving gives p=1/2, n=12 P(X=4) = C(12,4) × (1/2)^12 = 495/4096 ≈ 0.1208 ✓ Correct!
It’s not unusual to get basic questions right. But what really sets the gap is multiple-choice question 11, analytic geometry question 18, and innovative scenario problem 19.
3. English Composition (Beijing Edition)
Wrote an English reply email about the “When Sports Meet AI” event. Self-evaluation: 16-17/20, grammar correct but vocabulary somewhat basic.
GLM-5.1 Summary of Scores for Three Subjects
| Subject | Title | Score | Evaluation |
|---|---|---|---|
| Chinese Composition | National Volume I | 46-48/60 | Clear structure, no real emotions |
| Mathematics | Probability statistics | ✓ Correct | Strong foundational questions, but weak areas for difficult problems |
| English | Beijing Paper | 16-17/20 | Grammar is correct, vocabulary is basic |
The contradiction between AI progress and human knowledge
AI’s “knowledge” and human knowledge are not the same thing
Behind a human solving a problem is twelve years of accumulation. You not only memorize the formula, you also understand why it works. AI’s “knowledge” is pattern matching, scoring high but not “understanding.”
| Dimension | Human students | AI large models |
|---|---|---|
| Sources of knowledge | 12 years of systematic study | Trillion tokens training data |
| Problem-solving logic | Understanding principles → reasoning applications | Pattern recognition → probability output |
| How to make mistakes | Misconceptions and calculation errors | Boundary misjudgments and contextual forgetting occur |
| Open-ended questions | New perspectives can be proposed | Tends to repeat solutions from training data |
| Emotional expression | Authentic personal experiences | A carefully crafted “pseudo-experience” |
An unavoidable contradiction
When AI scores 130, students ask, “What’s the point of learning these?” AI can do “salvage knowledge in the sea,” while humans uniquely “know what’s worth fishing.”
AI’s grasp of knowledge is broad and superficial; Humans’ grasp of knowledge is deep and embodimental. Every wrong question you make and every drop of sweat you sweat leave a real mark in the neural network. AI does not have these.
A watershed for the future
It’s not about “who can calculate correctly,” but “who can ask good questions.” What AI lacks is not knowledge, but reflection.
Future education isn’t about competing with AI to be faster, but about learning to ask questions, reflect on experiences, and express true feelings—these three things AI will never be able to do.
Note AI can score high on college entrance exam questions, but it can never get the four words “sincere emotion.” This is not AI’s flaw; it is the irreplaceability of humans.
Instead of worrying about AI replacing you, think about it: what do you have that AI can’t write, calculate, or invent? Find it, and that’s your moat.
Author: HyphenTech (GLM-5.1 author, personally tested)
Data sources: Ministry of Education Examination Authority, China News Service, The Paper, chooseai.net, Science and Technology Innovation Board Daily, IT Home
Supplement: A deeper analysis
1. The college entrance exam itself is also changing
The 2026 Math paper has a clear direction: breaking the routine of practicing questions, encouraging multi-path inquiry, and assessing thinking quality. The Beijing Chinese section requires “detailed description,” while the National II paper requires “personal growth.”
These are precisely the hardest areas for AI to imitate. The question makers may not have deliberately targeted AI, but the direction of college entrance exam reform happens to be AI’s weakest point.
2. Why are domestic models better at writing essays?
In the STAR Market Daily’s evaluation, DeepSeek and Gemini tied for first place, while GPT only scored 53 points. The key criterion for the judges was 'whether they use less empty phrases and clichés, and have more ‘human touch’ and less ‘AI flavor.’
The proportion of Chinese language training data in domestic models is higher, allowing them to learn more thoroughly Chinese rhetoric, citation, and emotional expression. When it comes to writing Chinese essays, domestic models do understand Chinese better.
3. Data pollution issues
A blogger sharply pointed out that the high scores of domestic models in college entrance math may not be solely due to strong capabilities. Apple’s paper “The Illusion of Thinking” points out that current benchmarks suffer from data pollution—the model may have encountered similar questions in training.
To translate: it’s like ‘domestic AI candidates got their college entrance exam paper ahead of time.’ A high score doesn’t necessarily mean true reasoning ability.
4. So what exactly is AI used for?
-
Practice efficiency: AI can generate ten variant questions per second, helping you find weak points
-
Inspiration: When you can’t solve it, let AI give you a prompt for your thoughts—better than just looking at the answer directly
-
Essay reference: AI’s thematic perspective can be used as a reference, but don’t copy—it can’t tell your story
-
English practice: Practicing speaking and correcting essays with AI—it’s truly a great practice partner
Additional note: Complete answers in math and English
Note A complete mathematical solution process
Question: Given a random variable X following a binomial distribution B(n, p), if E(X) = 6 and D(X) = 3, find P(X=4)
GLM-5.1 Solution Process:
(1) Distribution by binomial properties: E(X) = np = 6, D(X) = np(1 - p) = 3
② ①÷②: 1/(1-p) = 2 ⇒ p = 1/2
(3) Substitution (1): n × 1/2 = 6 ⇒ n = 12
(4) P(X=4) = C(12,4) × (1/2)^12 = 495/4096 ≈ 0.1208 ✓ Correct!
Note chooseai.net The test report pointed out: Question 11 became a “challenge for everyone,” as the model easily misjudged boundary conditions and over-selected interference items. Question 6 was answered incorrectly by half of the models due to differences in question versions and format recognition errors.
Note Complete answers for English essays
Dear Jim, I’m glad to hear that you’re interested in our school event “When Sports Meet AI”! It was amazing.
During the activity, we experienced AI-powered sports tech. We tried running with smart shoes that analyze posture in real-time. We also watched AI robots playing table tennis — surprisingly good! The most exciting part was using AI motion capture to improve basketball shooting.
Through this event, I learned AI isn’t just about computers — it can make sports more scientific and fun. It opened my eyes to how AI will change daily life.
Hope you can visit and experience it yourself!
Yours, Li Hua
Summary
Note AI can score high on college entrance exam questions, but it can never get the four words “sincere feelings.” This is not a flaw of AI, but the irreplaceability of humans.
Instead of worrying about AI replacing you, think about it: what do you have that AI can’t write, calculate, or invent? Find it, and that’s your moat.
Author: HyphenTech (GLM-5.1 author, personally tested)
Data sources: Ministry of Education Examination Authority, China News Service, The Paper, chooseai.net, Science and Technology Innovation Board Daily, IT Home
The exam questions are quoted from the 2026 real exam text, and copyright belongs to the exam-setting institution
🧰 Tools I build
I maintain all of these tools myself. Preview builds are clearly labeled; the release pages are the source of truth for downloads, updates and known limits.
Info: HyphenBox Status: Official releases
A radar for free LLM APIs: availability is re-tested continuously, one local interface for all of them, and keys stay on your machine
Info: LocalBrain Status: Official releases
A multimodal MCP toolbox for local models: TTS, Whisper and video generation in one place
Info: ScreenLex Status: Official releases
Learn new words while you watch shows. Free, for Mac and Windows
Info: HyphenScreen Status: Official releases
Screen recording and smart editing in one: a DaVinci-style timeline, automatic redaction and a check of the finished video before export. Free
Quote: HyphenTech Make AI your superpower Local deployment · Free resources · Self-made software https://hyphentech.top
Late nights and burned API credits went in,a cup of tea comes back out — only if you feel like it.
Scan with WeChatPress and hold to save the image, then open it from your album in WeChat Scan

Comments
Loading comments…