Here's a question every standards-based grading teacher runs into sooner or later.
A student has five scores on the same standard this term. The first was rough. The last was great. The ones in between were okay. So... are they at Meets or not?
If you just average them, the rough start drags everything down. If you only look at the last score, one lucky day could carry the whole term. And if a parent asks, "How did you get that level?", you need an answer you can actually explain.
That's what this guide is about. No jargon, no formulas you need to memorize. Just the five methods, what each one does, and a worked example you can check with a calculator.
If you're new to standards-based grading itself, start with our complete guide to standards-based grading and come back. This post assumes you already have standards and scores, and you're trying to decide what they add up to.
What does "calculating mastery" actually mean?
In a traditional gradebook, you add up points and get a percentage for the class. In a standards-based grading system, you do something different. You look at each standard on its own, gather every piece of work that shows that skill, and decide what level the student has reached. That level is their mastery level for that standard.
The work you use is called evidence. A quiz question on fractions is evidence for the fractions standard. A paragraph where a student quotes the text is evidence for the "cite textual evidence" standard. Over a term, each student builds up a small pile of evidence for each standard.
"Calculating mastery" just means: how do you turn that pile into one level?
Most schools use a four-level mastery grading scale, something like Below, Approaching, Meets and Exceeds. Scores map onto those levels using cut-offs. In this guide we'll use a common set:
Your school's scale might use different words (Beginning, Developing, Proficient, Advanced) or numbers (a 1–4 scale). That's fine. The five methods work the same way on any standards-based grading scale.
Meet Maya and Leo
The easiest way to understand the five methods is to watch them work on real numbers. So let's meet two eighth graders. Both are working on the same reading standard: citing evidence from a text to support what they say about it.
Each of them has five pieces of evidence on that standard this term, in the order they happened.
Maya had a slow start. She spent most of the term at Approaching, then something clicked, and her unit essay was the best thing she wrote all term.
Leo was steady all the way through, solidly at Meets. Then his unit essay went badly. Maybe he was sick. Maybe he rushed it. Maybe he actually hadn't got it as well as it looked.
Keep these two in mind. We'll run both through every method, and you'll see how much the method matters.
Quick note: we're treating all five pieces of evidence as equally important for now. Later we'll look at weighting, where a unit test can count for more than an exit ticket.
Method 1: Simple average (every piece counts equally)
How it works: Add up all the scores and divide by how many there are. It's the method everyone already knows.
- Maya: (55 + 62 + 70 + 72 + 91) ÷ 5 = 70% → Approaching
- Leo: (82 + 85 + 80 + 84 + 58) ÷ 5 = 77.8% → Meets
When it works well: When every piece of evidence is a fair, finished measure of the same skill, and you don't expect the skill to grow much during the term. It's also easy to explain to anyone.
The catch: It punishes a slow start forever. Maya can read and cite evidence at an Exceeds level now, but her September exit ticket is still pulling her down in November. That's the main thing standards-based grading tries to fix. Grading expert Ken O'Connor put it simply in his book A Repair Kit for Grading: don't rely only on the average; look at other ways of summing up the evidence, and use your professional judgment.
Method 2: Decaying average (recent work counts more)
How it works: The newest piece of evidence gets most of the weight, and everything before it shares the rest. A common setting gives the newest score 65% of the weight and the average of all the older scores 35%.
Think of it like this: the latest thing a student did is the best clue to where they are now, but the earlier work still counts for something.
- Maya: newest = 91. Older average = (55 + 62 + 70 + 72) ÷ 4 = 64.75. So (91 × 0.65) + (64.75 × 0.35) = 59.15 + 22.66 = 81.8% → Meets
- Leo: newest = 58. Older average = (82 + 85 + 80 + 84) ÷ 4 = 82.75. So (58 × 0.65) + (82.75 × 0.35) = 37.7 + 28.96 = 66.7% → Approaching
When it works well: For skills that build over time, like writing, reading and problem solving. It rewards growth, which is why it's one of the most popular choices in standards-based grading. You can usually adjust the weight too. Make the newest piece count more (say 80%) and the level reacts faster to new work.
The catch: Look at Leo. One bad essay at the end knocked him from Meets down to Approaching, even though four out of five pieces of his work said Meets. The decaying average is only fair if the newest piece is a solid piece of evidence. If it isn't, you'll want to step in (more on that below).
Method 3: Most recent attempts (only the last few count)
How it works: Ignore the older work and average only the last few pieces of evidence. Three is a common number.
- Maya: (70 + 72 + 91) ÷ 3 = 77.7% → Meets
- Leo: (80 + 84 + 58) ÷ 3 = 74% → Approaching (by just one point)
When it works well: When you want the level to show what a student can do right now, and older work really doesn't matter anymore. It's a good fit for courses where students keep practicing the same standard and you reassess often. O'Connor makes the same point in his grading fixes: when learning grows over time, give more weight to recent achievement.
The catch: It throws away evidence. If a student had two strong months and one bad week, only the bad week might be left in the window. It also needs you to collect plenty of evidence, or "the last three" is basically the whole term anyway.
Method 4: Most frequent level (the level they hit most often)
How it works: Turn each score into a level first, then pick the level that shows up most. (Math teachers will know this as the mode.) If there's a tie, many systems round up to the higher level.
- Maya: Below once, Approaching three times, Exceeds once. Most frequent = Approaching
- Leo: Meets four times, Below once. Most frequent = Meets
When it works well: When you want a level that doesn't swing around because of one really good or really bad day. It's steady. For Leo, this is the method that best matches what most teachers would say about him: he's a Meets student who had one off day.
The catch: It's slow to notice real growth. Maya's breakthrough essay barely registers, because she only hit Exceeds once. It also works badly with very little evidence. With three pieces of work that are all different levels, there isn't really a "most frequent" at all.
Method 5: Highest level (their best work counts)
How it works: Take the student's best score and use that.
- Maya: best = 91 → Exceeds
- Leo: best = 85 → Meets
When it works well: When the question is simply "has this student shown they can do it, at least once?" It fits skills where showing it once really is proof, like a lab safety check or a specific procedure. Some teachers also like it because it's encouraging: your best work is what counts.
The catch: One good day can make it look like a student has mastered something they haven't. Maybe Maya's essay was her best writing ever, or maybe it was an easy prompt. With this method, you'd never know, because the other four pieces don't count at all.
The five methods side by side
Here's everything in one place. Same students, same five scores, five different ways of adding them up.
Maya ends up at three different levels depending on the method: Approaching, Meets or Exceeds. Leo lands at Meets or Approaching. Nobody did anything different. Only the maths changed.
That's the most important thing to take away from this post. Choosing a calculation method is a teaching decision, not a technical one. Each method is really a statement about what you believe. The decaying average says "students can grow out of a bad start." The simple average says "all of this counts." The highest level says "show me you can do it once." None of them is wrong. But you should pick one on purpose, not by accident.
Which method should you use?
If you want one answer: for most skills that build over time, a decaying average is a good default. It rewards growth, still remembers earlier work, and is easy to explain in one sentence. That's why so many standards-based grading systems use it.
But the better answer is to ask yourself three questions.
- Does this skill grow over the term? Writing, reading, reasoning and problem solving usually do. Use a method that leans on recent work: decaying average or most recent attempts. For knowledge that's either learned or not, like a set of vocabulary or a lab procedure, a simple average or the highest level can work fine.
- How often do students get another chance? If you reassess the same standard often, "most recent attempts" works well. If a standard only comes up two or three times a term, stick with a method that uses all of it, like a decaying or simple average.
- How do you want to treat a bad day? If one bad piece shouldn't move a steady student, the most frequent level protects them. If you'd rather catch a slide early, a recency method will show it.
Whatever you choose, use the same method for the whole class, write it down, and tell students and families. A method you can explain in one sentence is worth more than a clever one nobody understands.
Three settings that matter as much as the method
The method gets all the attention, but three other choices change the result just as much.
1. The scale and the cut-offs. The same 74% is Approaching if Meets starts at 75, and Meets if it starts at 70. Leo's "most recent" result sits exactly on that line. Decide your cut-offs as a department, before the term starts, and keep them. (We'll cover scales like Beginning–Advanced and Marzano's 1–4 scale in more detail in a separate guide.)
2. Weighting. Not all evidence should count the same. Most standards-based grading teachers agree that practice work shouldn't count toward mastery at all. It's for learning, not for grading. O'Connor's fixes say the same: use summative evidence to decide a grade, not practice. A common setup is to let practice count zero while still showing it in the record, and to let a unit test count more than a quiz. Remember Leo: if his unit essay counts double, even the simple average drops him to (82 + 85 + 80 + 84 + 58 + 58) ÷ 6 = 74.5%, which is Approaching. That might be right, or it might be a sign to look closer.
3. The minimum amount of evidence. Don't give a level too early. A level based on one quiz isn't a level. It's a guess. Many teachers wait until there are at least three pieces of evidence, and show "insufficient evidence" until then. O'Connor recommends exactly this kind of "I" for insufficient evidence instead of guessing or putting in a zero. How many pieces is "enough" is a big question, and we'll give it its own guide soon.
And one thing that matters most of all: you can always overrule the maths. If Leo's essay was a one-off and you've since seen him do it well in class, record that. A good system lets you set the level yourself, with a short reason, without deleting the evidence underneath. The calculation is there to help your judgment, not to replace it.
Five mistakes to avoid
- Switching methods halfway through the term. Every student's level can change overnight, and families will notice. Pick a method before the term starts and stick with it.
- Counting practice work. Homework and first drafts are for learning. If they count toward mastery, students learn that mistakes are expensive, which is the opposite of what you want.
- Giving a level too early. One piece of evidence is not enough to say where a student is. Wait until you have a few.
- Putting in zeros for missing work. A zero isn't evidence of what a student knows. It's evidence that something's missing. Mark it as incomplete and get the evidence instead.
- Letting the maths make the decision. The calculation is a starting point. If you know something the numbers don't, like a one-to-one reassessment or a bad week at home, you're allowed to act on it. Just write down why.
Doing this without a spreadsheet
You can do everything in this guide with a spreadsheet and some patience. Plenty of teachers do. The hard part isn't the maths. It's keeping it up for 120 students and 30 standards across a whole term, and then being able to show a parent the work behind each level.
That's what the standards mastery gradebook in GradingPal is built for. Here's how it handles what we've covered:
- All five methods, your choice. Simple average, decaying average, most recent attempts, most frequent level and highest level. You set it per class. For the decaying average, the newest piece counts 65% by default, and you can change it anywhere from 50% to 90%. For most recent attempts, you pick how many (3 by default).
- Your scale, your cut-offs. Below, Approaching, Meets, Exceeds; Beginning, Developing, Proficient, Advanced; or a Marzano-style 1.0–4.0 scale with half points. Edit the cut-offs to match your school.
- Weighting built in. By default, summative work counts and formative work doesn't. You can set it so a test counts twice as much as a quiz, and practice stays visible without moving the level.
- No early guesses. Until there's enough evidence, the cell says "insufficient evidence (2 of 3)" instead of a level. You choose the minimum.
- The maths, shown. Under every student's result, one line tells you which method was used and how many pieces of evidence went into it. Click any point and the actual graded work opens.
- You overrule it any time. Set a level yourself with a reason like "re-assessed one-on-one after the exam". It shows as "set by teacher", the evidence stays, and you can switch back.
- No AI in the maths. AI helps suggest which standards a question or rubric line matches, and you approve every one. But the mastery calculation itself is plain arithmetic, the same every time.
The evidence comes from grading you're already doing, across essays, worksheets, quizzes, exams and lab reports, typed or handwritten. You can also turn the levels into standards-based report cards for families, written in plain language, with each student's recent work listed underneath.
The bottom line
The same five scores can make a student look like three different learners. That's not a flaw in standards-based grading. It's a reminder that the method you choose says something about what you value.
So choose it on purpose. Lean on recent work for skills that grow. Keep practice out. Wait for enough evidence. Show your working. And when you know better than the maths, say so.
Want to try the methods on your own class without building a spreadsheet? Start free with GradingPal and turn on the standards mastery gradebook for one class. You can switch between the five methods and watch what changes.