GradingPal 2.0 is here — see everything that's new

AI essay grading research: inspect the scores, feedback, and consistency

GradingPal's published benchmark examines how selected language models score typed English essays, assess writing traits, generate feedback, and respond to different prompting and consistency tests. Explore the results alongside their methods and downloadable source material.

Published report: December 2025 · Five experiments · Essay assessment

Read each result with the task that produced it

The report examines several different questions. An overall essay score, a score for one writing trait, a model-judged feedback rating, and repeatability across runs measure different things. The tables and charts belong to their specific experiment and configuration.

This benchmark does not establish a single accuracy percentage for the GradingPal product. It does not measure handwritten math, science calculations, oral assessment, or student attainment. Those questions require evidence from the relevant task and setting.

Five questions about AI assessment

How closely do overall essay scores match the reference scores?

Experiment 1 compares model scores with the essay dataset's reference scores, using agreement and error measures.

How well do models assess individual writing traits?

Experiment 2 examines cohesion, syntax, vocabulary, phraseology, grammar, and conventions in a separate multi-trait task.

How is generated feedback rated?

Experiment 3 uses a panel of language models to rate feedback for specificity, actionability, accuracy, tone, and pedagogical value.

How stable are scores across repeated or changed inputs?

Experiment 4 examines repeated grading, paraphrased essays, and different temperature settings.

How do prompting strategies affect the results?

Experiment 5 compares rubric-only prompts, prompts with scored examples, step-by-step prompting, and the tested native reasoning configurations.

Understand the measures before comparing models

Quadratic weighted kappa (QWK) measures agreement between ordered ratings while accounting for chance agreement and weighting larger disagreements more heavily. It is not the percentage of essays graded correctly. A QWK of 0.80 does not mean 80% exact accuracy.

Exact match is the percentage of evaluated responses for which the model score matches the reference score exactly. Within one point is the percentage within one score point of the reference; interpret that distance in the context of the task's scoring scale.

Mean absolute error (MAE) is the average absolute distance between the model score and the reference score. Lower is better on the same task and scale. Other reported measures, including RMSE and correlation, provide different views of the same evaluation and should retain their labels.

For the metric definitions, see the scikit-learn documentation for Cohen's kappa and quadratic weighting and mean absolute error.

Feedback ratings in this report are judgments produced by the specified model panel. They are not student learning results or ratings from a teacher panel. Consistency concerns stability across the tested runs; a stable score can still disagree with the reference score.

Experiment 1: Compare overall essay scores with reference scores

The holistic experiment assesses one overall score per essay. The published summary includes six model rows, with different numbers of evaluated responses. Compare the agreement and error measures together rather than selecting a model from one headline number.

Published model labelEvaluated responsesQWKMAEExact match
Grok 4.14040.6890.71537.9%
Gemini 3 Pro2500.6800.61443.2%
Claude 4.52330.6650.58848.1%
Gemini 3 Flash2500.5900.68744.2%
GPT-5.12500.5210.89628.8%
GPT-5.1 Mini2500.3290.96829.6%

These are rounded values from the published experiment summary. The unequal evaluated counts mean the rows should not be described as a comparison on an identical complete set without checking the underlying records. A model with stronger QWK does not necessarily have the highest exact-match rate or lowest error.

What an educator can examine: Where does the proposed score differ from the reference, and what does the rubric say about the disputed evidence? A model ranking does not replace that review of the response.

Holistic Scoring Performance
Comparing models across key metrics

Charts scroll horizontally on small screens. The table above provides the reported values.

Complete archived summary, including evaluated counts
ModelN SamplesQWKMAERMSEPearson rExact Match %Adjacent ±1 %
Grok 4.14040.6890.7150.9560.77237.87191.089
Gemini 3 Pro2500.6800.6140.8410.76443.20094.400
Claude 4.5 Sonnet2330.6650.5880.8570.75748.06993.562
Gemini 3 Flash2500.5900.6870.9610.67544.20687.983
GPT-5.12500.5210.8961.1310.70028.80082.400
GPT-5.1 Mini2500.3290.9681.2360.53329.60075.200

Source: data_package/experiment_1_holistic/analysis/exp1_holistic_summary.csv. Display values rounded to three decimal places; the download retains original precision. Exact-match and adjacent-score columns are percentages.

Selected examples from the published report

11th Grade

High Performer with Strong Voice

Distance Learning Benefits

Human Score
6/6

"I hate coming to school" is something that I hear from my friends very often. Most teenagers go to a high school. Some find it very difficult and are searching for an alternative... Although some may believe otherwise, students would benefit from attending classes from home because school can be extremely stressful for students with poor mental health...

Grok 4.1 logoGrok 4.1
5/6
-1
Claude 4.5 logoClaude 4.5
4/6
-2
Gemini 3 Flash logoGemini 3 Flash
4/6
-2
Insight

Compare the scoring explanations with the rubric and the response.

Review question

Compare the scoring explanations with the rubric and the response.

Experiment 2: Examine writing traits separately

The multi-trait experiment evaluates cohesion, syntax, vocabulary, phraseology, grammar, and conventions. The archived summary reports results for six models. The trait-level detail shows why a single overall result can conceal differences between the kinds of judgment being requested.

Read each trait's result alongside the scoring task and available sample. This experiment uses a different task and dataset context from the holistic experiment. Comparing their headline QWK values alone does not isolate the effect of splitting a rubric into traits.

What an educator can examine: Check whether a comment about grammar, style, or organization is supported by the submitted writing and relevant to the criterion being assessed.

Writing-trait detail from the published report

TraitDescriptionReported QWK
CohesionHow ideas connect and flow0.529
ConventionsSpelling, punctuation, capitalization0.435
VocabularyWord choice appropriateness0.4
GrammarUsage and agreement0.39
SyntaxSentence structure0.363
PhraseologyPhrases, idioms, collocations0.311
Complete archived summary, including evaluated counts
ModelN SamplesQWK (avg)QWK (std)MAE (avg)Pearson r (avg)
Gemini 3 Flash4990.3700.1060.5810.634
GPT-5.14990.3480.0860.5590.627
GPT-5.1 Mini4990.3080.0350.5360.446
Grok 4.14990.2640.0460.6090.484
Gemini 3 Pro4990.1210.0650.8260.603
Claude 4.5 Sonnet4990.0890.0320.9290.574

Source: data_package/experiment_2_multi_trait/analysis/exp2_multi_trait_summary.csv. Display values rounded to three decimal places; the download retains original precision. Exact-match and adjacent-score columns are percentages.

Selected writing-trait examples

Easiest for AI

Cohesion Assessment

AI
4.5
Human
4.5

First of all, students with poor mental health would greatly benefit... Furthermore, online classes provide flexibility... In conclusion, distance learning offers many advantages.

AI Assessment

Strong transitions ('First of all', 'Furthermore', 'In conclusion') create clear logical flow between paragraphs.

Human Assessment

Excellent paragraph structure with consistent topic sentences and smooth transitions.

Review the writing-trait evidence

Check the evidence for each writing-trait judgment.

Experiment 3: Inspect the ratings for AI-generated feedback

Five models generated essay feedback, which was evaluated by three model judges across five dimensions. The reported overall ratings range from 4.57 to 4.71 out of 5. The dimension scores give more detail than the overall average.

Specificity asks whether the comment refers to the student's work. Actionability asks whether it gives a concrete improvement step. Accuracy concerns whether it describes the essay correctly. Tone concerns the way the feedback addresses the learner. Pedagogical value asks whether the comment helps explain the writing principle involved.

These are model-judged ratings under the experiment's conditions. They do not show how students understood the comments, whether teachers endorsed them, or whether the feedback improved later work.

What an educator can examine: Can the student locate the issue, understand why it matters, and take a useful next step? Review all three rather than relying on a supportive tone alone.

Model-judged feedback ratings
Feedback Quality Scores
Evaluated by panel of 3 LLM judges across 5 dimensions (scale: 1-5)
GPT-5.1 logoGPT-5.1
Gemini 3 Flash logoGemini 3 Flash
Claude 4.5 logoClaude 4.5

The numeric table below provides the same reported measures. Charts scroll horizontally on small screens.

Feedback ratings by dimension (out of 5)

ModelOverallSpecificityActionabilityAccuracyTonePedagogical value
GPT-5.14.714.964.74.854.964.09
Gemini 3 Flash4.694.914.784.774.644.33
GPT-5.1 Mini4.654.994.814.854.534.06
Claude 4.54.594.874.564.894.524.09
Grok 4.14.574.94.484.434.944.08

Selected feedback examples

Published report example

Gemini 3 Flash logoGemini's Teaching Approach

Strengths

You have gathered relevant evidence from various global contexts, such as the statistics regarding Vauban, Germany, and the smog levels in Paris.

Areas for Improvement

Your essay lacks a clear thesis statement. In the first paragraph, you ask questions but do not state a firm position. Try: "While cars offer convenience, cities should transition to car-free models to combat pollution."

Next Steps

1. Proofread for spelling errors like "gicing" and "partical." 2. Add a clear thesis statement at the end of your introduction.

Provides model sentences to emulateExplains WHY suggestions matterTeaches writing concepts
Compare the published feedback excerpts

Gemini 3 Flash

Strengths
You have gathered relevant evidence from various global contexts, such as the statistics regarding Vauban, Germany, and the smog levels in Paris.
Improvements
Your essay lacks a clear thesis statement. In the first paragraph, you ask questions but do not state a firm position. Try: "While cars offer convenience, cities should transition to car-free models to combat pollution."
Next steps
1. Proofread for spelling errors like "gicing" and "partical." 2. Add a clear thesis statement at the end of your introduction.

GPT-5.1

Strengths
You clearly share your opinion, especially in sentences like, "Thats not smart," which shows a strong stance. You also use examples from different places, which is great evidence to support your argument.
Improvements
Your introduction needs a clearer claim. For example, instead of "Who wouldnt want to have a car?" you could write: "Car-free cities are a bad idea because people need cars to get to work and school on time."
Next steps
1) Write a clear thesis sentence. 2) Combine short, choppy sentences into longer ones. 3) Revise for spelling and punctuation.

Grok 4.1

Strengths
Your essay has a clear stance against car-free cities, and you use multiple sources to support your views. The reference to Paris's smog issues is particularly effective.
Improvements
Focus on organizing your body paragraphs better. Each paragraph should have one main idea with supporting evidence. Avoid jumping between different topics.
Next steps
1. Create an outline before writing. 2. Use topic sentences for each paragraph. 3. Check for run-on sentences.

Experiment 4: Test score stability across repeated runs

The consistency experiment contains three distinct tests. Keep their results separate: repeatability on the same essay is different from sensitivity to changed wording or sampling settings.

Repeated grading of the same essay

The published test grades the same essay repeatedly and compares the resulting scores. Some tested configurations produced identical scores across the observed repetitions. That is evidence about these runs, not a guarantee of identical results for all future inputs or model versions.

Original and paraphrased essays

The semantic-equivalence test compares scores for an original essay and a paraphrased version. Review the size of any difference alongside the actual text; calling two responses equivalent does not itself establish that every rubric-relevant feature is unchanged.

Different temperature settings

The temperature test reports results at 0.0, 0.5, and 1.0. The pattern varies by model; the published table does not show a strictly increasing variance at every step for every model. Inspect the curve rather than treating temperature as a guarantee of a particular score.

What an educator can examine: Review both correctness and stability. An assessment can repeat the same incorrect judgment, and a change in score deserves inspection against the rubric.

Repeated grading — identical resubmission

The numeric table below provides the same reported measures. Charts scroll horizontally on small screens.

Identical-resubmission results

ModelExact match (%)Variance
Gemini Pro1000
Gemini Flash1000
Claude 4.591.70.015
GPT-5.179.20.035
Grok 4.162.50.075
GPT-5.1 Mini500.09
Original and paraphrased essays

The numeric table below provides the same reported measures. Charts scroll horizontally on small screens.

Semantic-equivalence results

ModelExact match (%)Mean score difference
Gemini Flash93.80.062
Gemini Pro89.60.104
GPT-5.1 Mini87.50.125
Grok 4.185.40.146
Claude 4.583.30.167
GPT-5.1750.25
Score variance at different temperature settings

The numeric table below provides the same reported measures. Charts scroll horizontally on small screens.

Reported variance by temperature

TemperatureClaudeGemini FlashGemini ProGPT-5.1Grok
0.00.011000.0510.081
0.50.0280.0290.020.0310.13
1.00.0410.0390.0790.0820.161
Reported composite consistency ranking

This is the report's composite of consistency measures, not grading accuracy. The displayed composite averages the three reported component scores. The temperature component is a transformed score, not raw variance; use the archived analysis to inspect its construction and missing-run handling.

Published composite and component values

ModelIdenticalSemanticTemperature componentComposite
Gemini Flash10093.880.491.4
Claude 4.591.783.379.784.9
Gemini Pro10089.660.783.4
GPT-5.1 Mini5087.510079.2
GPT-5.179.27559.271.2
Grok 4.162.585.419.855.9

Selected consistency examples

Identical Resubmission

Gemini 3 Flash logoRepeated scores in the published example

Same essay submitted 5 times, same score every time

Published report example

"The advantages of limiting car usage are numerous. In Vauban, Germany, residents have given up their cars..."

5 Submissions:33333

Inspect what changed between the compared inputs and scores.

Experiment 5: Compare rubric prompts and scored examples

The prompting experiment compares four strategies: zero-shot rubric prompts, few-shot prompts with scored examples, chain-of-thought prompts, and the native reasoning configurations included in the evaluation. Not every model has a result for every strategy.

The reported few-shot QWK is higher than the reported zero-shot QWK for each of the six models in that comparison. The size of the difference varies. For example, the displayed Claude values move from 0.705 to 0.793, while the GPT-5.1 values move from 0.718 to 0.722.

These are changes in an agreement metric for the tested configurations. They should not be rewritten as a universal percentage improvement in grading accuracy, or as a guarantee that adding examples will improve every assessment.

What an educator can examine: Do the criteria and examples communicate the intended scoring standard clearly? Evaluate the resulting assessments rather than assuming that a more elaborate prompt is better.

Prompt strategy comparison (QWK)

The numeric table below provides the same reported measures. Charts scroll horizontally on small screens.

Reported agreement by strategy

ModelZero-shotFew-shotChain-of-thoughtReasoning mode
Gemini Pro0.7390.80.7670.758
Grok 4.10.730.7970.724Not reported
Claude 4.50.7050.7930.780.427
GPT-5.10.7180.7220.6490.533
Gemini Flash0.6410.6660.6420.58
GPT-5.1 Mini0.50.5180.4820.385
Change in QWK: few-shot minus zero-shot

The numeric table below provides the same reported measures. Charts scroll horizontally on small screens.

Absolute change in QWK

ModelChange in QWK
Claude 4.50.088
Grok 4.10.067
Gemini Pro0.061
Gemini Flash0.025
GPT-5.1 Mini0.018
GPT-5.10.004

Few-shot results

ModelQWKExact match (%)Within one point (%)
Gemini Pro0.833.382.5
Grok 4.10.79731.587.1
Claude 4.50.79332.586.7
GPT-5.10.7221582.5
Gemini Flash0.66623.668.9
GPT-5.1 Mini0.51816.759.2
Complete archived summary, including evaluated counts
StrategyModelQWKExact_MatchAdjacent_1MAEN_Samples
Zero-Shotgpt-5.10.71816.12983.0651.024124
Zero-Shotgpt-5.1-mini0.50016.93558.0651.266124
Zero-Shotclaude-4-5-sonnet0.70530.36683.2460.869191
Zero-Shotgemini-3-pro-preview0.73929.03279.8390.889124
Zero-Shotgemini-3-flash-preview0.64125.51073.4691.01098
Zero-Shotgrok-4.10.73029.03280.6450.911124
Few-Shotgpt-5.10.72215.00082.5001.012120
Few-Shotgpt-5.1-mini0.51816.66759.1671.258120
Few-Shotclaude-4-5-sonnet0.79332.50086.6670.807120
Few-Shotgemini-3-pro-preview0.80033.33382.5000.806120
Few-Shotgemini-3-flash-preview0.66623.58568.8681.002106
Few-Shotgrok-4.10.79731.45287.0970.815124
Chain-of-Thoughtgpt-5.10.64917.50072.5001.108120
Chain-of-Thoughtgpt-5.1-mini0.48215.83355.0001.300120
Chain-of-Thoughtclaude-4-5-sonnet0.78031.66785.8330.833120
Chain-of-Thoughtgemini-3-pro-preview0.76732.50084.1670.842120
Chain-of-Thoughtgemini-3-flash-preview0.64225.00070.8331.050120
Chain-of-Thoughtgrok-4.10.72427.41980.6450.927124
Reasoning Modegemini-3-pro-preview0.75846.49196.4910.570114
Reasoning Modegemini-3-flash-preview0.58030.63183.7840.862111
Reasoning Modegpt-5.10.53328.94779.8250.912114
Reasoning Modegpt-5.1-mini0.38521.05369.2981.105114
Reasoning Modeclaude-4-5-sonnet0.42735.00086.5000.795200

Source: data_package/experiment_5_prompting/analysis/exp5_complete_metrics.csv. Display values rounded to three decimal places; the download retains original precision. Exact-match and adjacent-score columns are percentages.

Selected prompting examples

Zero-Shot vs Few-Shot

Claude 4.5 logoWith and without scored examples

Compare the published scores under the stated prompting strategies.

Published report example
Zero-Shot
Score: 2
Few-Shot
Score: 3
Human Score
Score: 3

Compare how the scoring standard is expressed in each prompt.

Use research to ask better questions about a grading tool

When evaluating AI assessment, ask which tasks were tested, which reference scores were used, how many responses were evaluated, and how disagreements were examined. Ask separately about the quality of feedback and the usefulness of the review workflow.

For your classroom, use assignments and rubrics you understand well. Include different response types and quality levels. Inspect where the proposed score differs from your judgment, whether the explanation identifies the right evidence, and whether the feedback gives the student a manageable next step.

For a school, combine this assessment review with curriculum fit, teacher usability, integration requirements, and privacy documentation. The benchmark can inform the questions; it does not complete a school's evaluation.

See how GradingPal supports assessment review

GradingPal provides rubric-based scores with explanations, editable feedback, classroom analytics, and standards mastery tracking. Teachers can examine the submitted work and change the assessment before deciding what to return.

Those product capabilities explain how a teacher can inspect and use assessment evidence. They are separate from the specific model configurations in this historical essay benchmark. See the AI grading workflow and read the responsible AI approach.

The GradingPal review workflow. Product context, separate from the essay benchmark.
The GradingPal review workflow. Product context, separate from the essay benchmark.

Download the source material and inspect the method

The published data archive includes input datasets, prompts, experiment outputs, and analysis files. It contains PERSUADE and ELLIPSE input files; the number of available input rows is different from the number of successfully evaluated responses in any particular experiment.

The included PERSUADE input CSV has 1,000 data rows and the ELLIPSE input CSV has 499. Some README descriptions, model labels, and dates differ from the files they describe. Read the experiment-specific files alongside the summary and use the reported result counts when interpreting a table. These source discrepancies should not be treated as a new or larger completed evaluation.

The report is presented as December 2025; parts of the archive contain different dates. The archive is provided in its published form for inspection. The README describes release for research purposes and refers to the underlying datasets' original licenses; that is not a blanket license for unrestricted reuse.

Suggested citation: GradingPal Research Team. Evaluating Large Language Models for Essay Scoring and Feedback: An Empirical Comparison. Published data package, version 1.0, report presented as December 2025. GradingPal. Include the archive link and your access date.

Treat this as a dated benchmark

The report has not been re-run in the materials published here. Changes in models, prompts, datasets, or product workflows require a new evaluation before making a claim about current performance. These results are neither a guaranteed floor nor a ceiling for later systems.

Questions about the research

Does this measure GradingPal's accuracy across every assignment type?

No. It examines the typed English essay tasks and model configurations described in the report. The results should be read with the corresponding methods and evaluated counts.

Does it establish handwriting, math, science, or oral-grading accuracy?

No. Those workflows need their own relevant evaluation. GradingPal supports additional assessment types, but this essay benchmark does not supply their accuracy figures.

Does it show improved student attainment?

No. Scoring agreement, model-judged feedback ratings, and consistency are different from evidence of classroom learning gains. A study of student progress would need to measure those outcomes directly.

How should I evaluate GradingPal for my teaching?

Use a familiar assignment and rubric. Review the scores, explanations, and feedback, including responses where your judgment differs. Consider whether the workflow makes assessment easier and whether the comments help students act. Explore AI grading.

When was the research published, and is it a current benchmark?

The report is presented as December 2025, with some inconsistent dates in the archived material. No re-run is documented here. Treat the results as historical evidence for the specified tasks and configurations, not as measurements of today's models or product.

Explore the evidence and the product

Inspect the data, review the assessment workflow, or talk with us about evaluating GradingPal in your setting.