College Application Season: Qwen AI Beats Human Consultants in Head-to-Head Test
As 12.91 million Chinese students face their college application decisions, a new benchmark from Yousong Lab pitted AI against 53 experienced human consultants. Qwen's college application agent scored 100% on 44 policy knowledge questions (humans averaged 89.3%), recommended 6 viable choices vs 5.3 for humans, and won 58 out of 100 anonymous preference tests.
💡 What You Will Learn
As 12.91 million Chinese students face their college application decisions, a new benchmark from Yousong Lab pitted AI against 53 experienced human consultants. Qwen's college application agent scored
📜 Table of Contents
44 Policy Questions: Qwen Gets All Right, Humans Average 5 Wrong
Yousong Lab designed 44 professional questions covering Gaokao policies, new-subject-selection requirements, parallel-volunteer rules, and special program policies. Qwen's Gaokao volunteer agent answered all 44 correctly — 100% accuracy. The 53 human consultants, averaging 4.6 years of experience, averaged 89.3% — roughly 1 wrong per 10 questions.
More importantly: when human consultants worked with Qwen's assistance, their accuracy improved AND their time spent dropped by ~27%.
Simulated Application: AI Hits 6 vs 5.3 for Humans
In a simulated 10-choice college application scenario: - Human consultants: average 5.3 viable admissions - Qwen AI plan: average 6 viable admissions, with 0 preference violations
The 0 preference violations is the key metric. College applications come with hard constraints: "I won't go to the Northeast," "I don't want to study medicine." Consultants sometimes push "better value" options against these constraints. Qwen was trained with nearly 400,000 simulated question patterns — including students changing their minds mid-session — so it strictly respects every hard constraint.
100 Anonymous Blind Tests: Experts Prefer AI Answers
Most revealing: 100 fully anonymous comparisons (assessors didn't know which was AI, which was human). Experts chose Qwen's answer 58 times, commenting it was "more stable in professional path analysis, risk warnings, and clarity of expression."
Context: the 53 human consultants average 4.6 years of experience — mostly mid-level practitioners. Qwen's backing is 8 years of Kuaishou Gaokao data — ~3,000 schools, 2,000+ majors with historical scores, rankings, trends, plus industry talent gap data and corporate hiring data. 4.6 years vs 8 years of data is an unfair comparison. But the reality is most families can only access mid-level consultants.
Three Technical Details
-
Adversarial AI student: Qwen trained a dedicated "AI student" model that simulates nearly 400,000 question patterns — sometimes changing preferences mid-conversation, sometimes changing topics. Through repeated adversarial training, Qwen learned to maintain consistent, high-quality output regardless of user behavior.
-
7 reward functions: Beyond correctness labeling by volunteer experts and adversarial training with the AI student, Qwen was optimized with 7 reward functions including response style and structural clarity. It's not just stacking knowledge — it knows "what language students and parents can understand."
-
Looking beyond Gaokao to careers: Qwen's knowledge base goes beyond scores and rankings. It incorporates industry talent gap data and corporate hiring needs, mapping them back to major choices. It doesn't just tell you "what you can apply for" — it tells you "where this decision leads in 3 years."
Personal Take
In a market of 10+ million test-takers with policies changing every year, quality college application consulting is extremely scarce. Good consultants charge 3,000-10,000+ RMB — out of reach for most families.
Qwen's real value isn't proving AI is smarter than humans. It's democratizing the average service quality — making it free or near-free for every student. On June 23 (score release day), students could enter their province, scores, and subject choices into the Qwen app for a personalized report.
One concern: as AI knowledge bases become more accurate, students' application strategies will converge. When millions of students use the same agent, the boundary between "hot" and "cold" majors may flatten rapidly, potentially causing wild score fluctuations in some programs within 2-3 years. This is a problem that will need addressing soon.
Written by our editorial team; tools listed here are tested or verified against public sources. Links point to official sites or GitHub repos for reference only — no paid placements.
