What Was I Really Grading?
Rigor, Equity, and Language Learning in the High School Spanish Classroom
By Erika Sanneman
I did not just set out to make grading easier and more functional for myself. I set out to make it more honest.
The Questions Underneath the Gradebook
I began teaching at a small private high school as one member of a Spanish department that ranged from approximately five to seven teachers over three years there. Before the school year began, I first taught remedial summer-school Spanish I and II for students who had not passed their courses and needed the credits for graduation. During my first regular school year, I taught Spanish I, Honors Spanish I, Spanish III, and Honors Spanish III.
I was also white-presenting in a predominantly Black school. Official demographic information —approximately three-quarters of the student body was Black and another portion was multiracial. Demographic forms, self-identification, and visual perception are not interchangeable, but the racial context of the school was unmistakable. My identity affected the authority I carried, the assumptions attached to me, and the way my decisions could be experienced. Grading never occurred in a social vacuum.
The school was actively examining racial equity, cultural relevance, inclusion, academic language, standardized testing, and the historical structures embedded in education. Faculty book clubs, professional development, department conversations, and my graduate coursework kept returning me to the same questions: What do we call rigorous? Is a practice rigorous because it produces deep learning, or because it is traditional? Which voices are treated as canonical? Who has learned the hidden rules of school before entering the room?
I was completing a master's degree in second-language education at the time. The questions I encountered in research courses were also alive in my school: how to teach what students genuinely needed while expanding whose knowledge, language, history, and excellence counted. I wanted students to understand the traditional academic references they would encounter in college and professional life without teaching the traditional canon as though it represented the whole world.
The central problem became both philosophical and practical: How could I create a grading system that was equitable, culturally informed, rigorous, transparent, and sustainable enough to survive an actual teacher's workload?
A Shared Textbook Was Not a Shared Course
Like teachers in most schools, I worked within established expectations. The department did not have a complete, teacher-written curriculum, but it had purchased textbooks, workbooks, online companions, and publisher-created assessments. We were expected to use them as the common framework for pacing and content.
Several teachers might teach the same course level, sometimes three or four of us at once. We tried to remain near the same chapter and to ensure that students had technically encountered the same vocabulary and grammar by the end of the year. Yet a shared textbook did not create identical learning experiences.
Teachers made different decisions about how long to spend on a topic, how often students spoke, how much culture and geography to include, which assignments were graded, how homework was weighted, and how frequently students were quizzed. One teacher's homework might be 10 percent of the grade and another's 30 percent. One might emphasize weekly quizzes while another relied on larger unit tests.
That variation raised a deceptively simple question: What did an A mean? Did it represent memorization, independent communication, assignment completion, improvement, compliance, or the ability to satisfy one particular teacher's preferences? Numbers looked objective, but the decisions underneath them were often deeply subjective.
The department did have boundaries. We were expected to use the textbook and workbook, and a portion of major tests came from the publisher's assessment bank. Within those boundaries, however, there was substantial freedom. I began using that freedom to ask whether each part of the grade measured the language skill I claimed it measured.
When Completion Looked Like Mastery
One of the first things I noticed was that completed practice did not always translate into independent mastery. The workbook and online platform offered many ways to check, reconstruct, or locate correct answers. I do not want to label all of that cheating. A student might have used a tutor, checked the textbook carefully, corrected the work until the platform accepted it, or received legitimate help.
What I could see was a mismatch. Some students earned nearly perfect homework grades because they completed everything diligently, yet they could not reproduce the same skill in conversation, writing, listening, or an independent classroom task. Their effort was real. The claim made by the grade was sometimes not.
This mattered especially in a college-preparatory private school. Students experienced intense pressure to earn honor-roll grades and protect transcripts. Families were investing money, time, and hope. It could become easy for everyone involved to assume that strong effort, tuition, and completed work should produce an excellent grade, and that an excellent grade necessarily proved excellent mastery.
Ideally, an A should communicate something meaningful about what a student knows and can do. But language learning is uneven. A student may understand far more than they can say. Another may speak fluidly but write inaccurately. A third may perform beautifully on controlled grammar exercises and struggle to understand natural speech.
Grades also carry imagined qualities they do not directly measure: intrinsic motivation, concentration, persistence, curiosity, organization, and the quality of a student's practice. Two students can submit the same worksheet after engaging with it in radically different ways. A percentage cannot fully contain that difference.
The problem was not that practice lacked value. The problem was treating practice and proof of learning as interchangeable.
Designing Assessments Around Language
I began balancing the types of questions on tests and quizzes. A student's cumulative grade should not depend too heavily on one paper-test format, especially when familiarity with multiple-choice, true-or-false questions and standardized testing may reflect previous schooling as much as language knowledge.
I tried to align each question with a discrete skill. If I wanted to assess vocabulary, the choices needed to test meaning. If I wanted to assess conjugation, the choices needed to test agreement. If I wanted to assess interpretation, students needed to make meaning rather than merely recognize a familiar-looking form.
I also stopped filling assessments with invented misspellings and malformed grammar unless error analysis was itself the goal. Research has found that exposure to incorrect spellings can affect later spelling of those same words, although this literature is not conclusive. However, it should not be treated as a universal rule.1 My practical solution was to use authentic distractors: every option might be a correctly formed Spanish word or conjugation, but only one was correct in that context.
Listening became part of every major assessment. Language is not only written. Students needed to interpret what they heard, notice key information, and respond to spoken Spanish. The structure gradually moved closer to the interpretive, interpersonal, and presentational modes emphasized in ACTFL's language-learning standards.2
I also paired major tests with cumulative projects. If summative work represented half of a quarter grade, for example, the test might count for 25 percent and the project for 25 percent. The test showed what students could demonstrate within a limited period. The project showed what they could build through sustained effort, feedback, rehearsal, revision, and preparation.
Projects included clear rubrics distributed several weeks in advance. Students knew whether they were being evaluated on comprehensibility, vocabulary, grammar, cultural evidence, interpretation, collaboration, or presentation. Transparency was not a courtesy added after the grading. It was part of the assessment design.
Access, Portability, and Universal Design
Students missed class for illness, appointments, athletics, performances, competitions, and school events. Some completed assessments in testing rooms supervised by adults who did not speak Spanish. That created a useful test for my own materials: Could another adult administer this assessment fairly without relying on my improvisation?
I made directions more explicit and self-contained. Audio needed to be prerecorded or consistently available. Procedures, time limits, and permitted resources needed to be clear. An assessment that only worked while I stood in the room explaining it was not equally accessible to students who had to take it elsewhere.
I also began turning individual accommodations into universal features whenever doing so did not alter the learning target. When one student needed larger print, I used that print size for everyone. I increased white space, simplified visual organization, made audio accessible, and created predictable directions. The student could still use a separate room when needed, but did not have to carry a visibly different test.
This approach aligned with Universal Design for Learning, which asks educators to anticipate learner variability and provide multiple means of engagement, representation, and action or expression.3 The objective was not to make every task easier. It was to remove barriers unrelated to Spanish.
Retakes as New Evidence
The most consequential change was allowing students to retake quizzes and tests within a defined window. Language develops through repetition, feedback, and renewed attempts. Treating one quiz as a permanent record of what a student would always know made little sense to me.
Students generally had seven school days from the original assessment to study, seek support, and attempt a new version. I eventually created five or six aligned versions of each quiz or test. The structures and learning goals remained the same, but the examples changed. Students could not simply memorize an answer sequence. They had to apply the skill again.
Iversionsquestions were incorrect but did not automatically supply every correction. Students reviewed the pattern of the errors: direct object pronouns, verb meaning, agreement, listening comprehension, vocabulary, or another clearly identified skill. The feedback shifted from “You got a 72” to “You understand the vocabulary, but you are not consistently matching the pronoun to the noun it replaces.”
The reassessment cycle was simple in principle: attempt, receive evidence, identify the gap, practice, and demonstrate again. Research and professional literature on reassessment support its potential to improve learning and reduce the high stakes attached to a single performance, while also warning that retakes need clear conditions and can become difficult to manage at scale.4
The first score was not erased because it had never mattered. It had shown where the student was then. The new score could matter because it showed where the student was now.
Revision Beyond Tests
Within another year, I realized that projects presented the same issue. Students might receive the rubric three or four weeks in advance and still misunderstand a requirement, lack a skill, or recognize only after seeing feedback what stronger work would look like.
I began allowing revisions to specific project components. A student who had not used the required vocabulary in an audio submission could rerecord that portion. A student who had missed cultural evidence could strengthen that section. The revised evidence replaced the part of the rubric it had changed.
Students generally had about a week after receiving the grade to review the rubric, seek assistance, and resubmit. If many students misunderstood the same requirement, I addressed it with the class. A widespread pattern was not automatically proof that everyone had failed to listen. It could also be evidence that my directions or instruction needed revision.
Support could not depend on a teenager asking for help perfectly. Some students were anxious about one-on-one conversations. Others were embarrassed, overscheduled, or unavailable during formal tutoring times. I met students in the minutes around bells, circulated quietly during class, allowed a friend to accompany them, or found less pressured locations such as the library.
The grade was intended to describe where a student was, not to pronounce what the student would always be capable of doing.
Resources Without Worshipping the Teacher
I created an online library of recorded lessons, slides, textbook references, examples, review materials, and links to other teachers explaining the same concepts. If another teacher's explanation made the idea click, that was a success. My goal was for students to learn, not to require that all learning pass through my personality.
I also tried to keep assessments transferable. Success should not depend on remembering my favorite color, a private classroom joke, or an incidental detail about me. Another qualified Spanish teacher should be able to recognize the skill being measured.
Across assignments, students spoke, listened, read, wrote, interpreted, presented, negotiated meaning, collaborated, and worked independently. No single weakness was allowed to consume the entire course grade. A student who struggled with public speaking still needed speaking practice, but one presentation did not erase all other evidence of language ability.
Group work was structured so that individual contributions could be identified. One student should not receive the same grade for work another student completed, and one anxious high-achiever should not be rewarded for quietly absorbing every responsibility. Collaboration was taught and assessed without allowing the group score to become a fog machine.
Why I Stopped Using Traditional Extra Credit
During my first year, I continued cultural extra-credit traditions used by previous teachers. They were often fun and gave me room to add topics beyond the required textbook. Over time, I recognized two problems.
First, presenting culture as extra credit implied that culture was extra: decorative, optional, and separate from the “real” academic work of grammar and vocabulary. Second, students could use unrelated points to compensate for weak language performance without addressing the missing skill.
I stopped offering extra credit that merely raised the number. Cultural and geographic knowledge became part of quizzes, tests, projects, classwork, homework, listening, and discussion. Students could improve a grade by improving the learning connected to that grade.
The route to a higher number became more evidence, not more points.
Privacy, Feedback, and the Work Worth Automating
I also stopped traditional routines such as exchanging papers, announcing scores aloud, and publicly collecting grades. Peer grading is not automatically prohibited by FERPA; the Supreme Court held that peer-graded papers are not education records before the teacher collects and records them.5 My decision was ethical and pedagogical rather than a claim that every such routine was unlawful.
Efficiency did not justify exposing a student's performance or creating avoidable embarrassment. Grades did not need to become public classroom currency.
I also stopped spending hours correcting every error when students rarely read the markings. I identified patterns and next steps without doing the entire intellectual task for them. Technology handled mechanical scoring through Google Forms, Sheets, Docs, the learning-management system, and the publisher's platform. Human attention was reserved for spoken communication, extended writing, comprehensibility, cultural interpretation, and the way skills worked together.
Automation handled arithmetic so I could spend more time on language.
Authentic Communication and Lexical Variation
As clerical grading became more efficient, classroom time became more dynamic. We staged conversations, practiced common exchanges, listened to real speakers, and watched videos created for audiences beyond language learners.
Students heard that Spanish was not one perfectly uniform system. Vocabularyan and pronunciation varied across countries, regions, generations, and communities. A word used in Colombia might differ from one used in Spain, Peru, Puerto Rico, Mexico, or the Dominican Republic.
Difference was not automatically error. It could be evidence of geography, history, migration, identity, and community. The textbook became one resource rather than the single gatekeeper of legitimate Spanish.
Open-Note Assessments and the Hidden Curriculum of Studying
My next change came from research on open-note and open-book assessment. The evidence is more nuanced than a simple claim that open notes never affect anyone except anxiou in minds students. Studies have found that students often perceive lower anxiety during open-note examinations, while performance effects depend on the course, exam design, preparation, and how students use their notes.6
With that caution, I began allowing students to use their own handwritten notes on nearly every assessment except direct vocabulary recall. They could not print my slides, use another student's notebook, or bring a commercially prepared guide. The resource had to be their own record of learning.
This exposed another hidden curriculum: schools expect students to take useful notes without consistently teaching them how. I embedded brief lessons on headings, organization, color coding, examples, tabs, note cards, retrieval, review, and rewriting confusing material.
I taught those strategies because they supported learning, but I did not grade notebook aesthetics. A beautifully color-coded notebook did not prove Spanish proficiency, and a chaotic notebook did not automatically disprove it.
The open-note policy created a natural consequence. Notes only helped when students understood and could apply them. Copying every word from a slide was useless if the student could not interpret what the words meant.
I did not collect formal quantitative data, and my classes varied too much for casual numerical comparisons to carry much weight. What I observed, however, was a substantial change in the number of students taking notes and in the care they gave those notes. Students noticed missing explanations, unfinished sentences, unclear examples, and gaps created by absences.
The notebook stopped being evidence of compliance. It became a tool.
The Mathematical Crater of a Zero
During these years, the school also adopted a minimum-grade policy. A completely missing submission received a 50, and any genuine attempt received at least the lowest F, approximately 64.
The policy responded to a real mathematical problem. On a conventional percentage scale, a zero occupies far more distance below passing than any other letter grade occupies above it. One missing assignment can create a numerical crater that later learning struggles to fill. Scholars such as Thomas Guskey have argued that the deeper problem is the use of percentage scales and averaging, not simply the zero itself.7
The new policy also created complications. Because my previous weighting system had not been designed around minimum grades, some averages appeared inflated. I had to revisit the balance among homework, practice, projects, tests, and demonstrations of mastery.
The challenge was to prevent one mistake from mathematically destroying a quarter without allowing the final grade to claim learning the student had not demonstrated. Minimum grades alone could not solve that. They needed to be paired with clear learning targets, meaningful reassessment, and summative evidence that mattered.
What Changed in the Classroom
The most visible change was reduced anxiety. Students still cared about grades, and families still worried about low scores, but one performance was no longer an irreversible verdict.
Conversations with parents changed as well. Instead of bargaining over points, I could say: Your child has not yet demonstrated this skill. Here is the evidence. Here are the resources. Here is the reassessment window. Here is how I can help. Here is what the student must do next.
Parents of struggling students and students with disabilities, anxiety, attention differences, processing needs, chronic absences, or intense schedules played an important role in shaping the system. Their questions exposed barriers I had not noticed. They knew how their children responded to pressure and which supports increased or diminished independence.
That did not mean every requested change was appropriate. It meant their concerns contained information. A family questioning a grade gave me an opportunity to test whether I could defend the relationship among the number, the evidence, and the stated learning goal.
Grace was not the opposite of rigor. A student might be sick, exhausted, grieving, returning from a competition, or simply experiencing an ordinary adolescent failure of planning. Reassessment helped distinguish a temporary barrier from a genuine gap in Spanish. If the student still could not demonstrate the skill after support and practice, the grade continued to show that gap. If the student learned it, the grade could change.
A System Built Through Collaboration
Although I was responsible for the final choices in my classroom, I did not create this system alone. My thinking developed through years of conversations with Cassandra Mahoney, Charles Shryock, Matt Buckley, Adam Greer, and Debbie Brennan.
They helped me interrogate different parts of the puzzle: cultural inclusion, grading, curriculum, assessment, educational theory, philosophy, and the rhythms required to make an idea work during a real school day. Their questions showed me where my thinking was incomplete. Their experiences gave me alternatives I might not have developed independently.
The system grew through department meetings, shared resources, disagreements, conversations between classes, questions after professional development, and the slow professional process of letting other people find weak seams in an idea before those seams hardened into practice.
Equitable teaching is often narrated as the work of one unusually reflective teacher. My experience was much more communal.
Transparency as an Equity Practice
From the beginning of each year, students and families knew the major assessment dates, category weights, retake rules, project revision windows, available resources, and the reasons behind the system. Rubrics and learning goals were visible before students began.
This did not prevent every misunderstanding. It reduced the power of hidden academic rules. When expectations remain implicit, students who already know how to decode teachers, organize long-term work, advocate with adults, or obtain outside help begin with an advantage.
Making the route explicit did not guarantee equal outcomes. It made success less dependent on already knowing how school works.
Transparency also protected flexibility from becoming favoritism. Everyone knew what opportunities existed, how to use them, and where the boundaries were.
When the System Met Scale
This system developed over approximately five years, but many of its most detailed features were refined during a three-year period when I taught AP Spanish Language, AP Spanish Literature and Culture, Women's Studies, Economics, Spanish III, and Spanish I. Several of those courses were small.
Small classes gave me time to provide complex feedback and regrade the same project repeatedly. Generous planning periods helped. I staggered major assignments so that the most labor-intensive work did not arrive all at once. During those years, I generally completed my work during contracted hours.
My final high-school year was different. I taught primarily Spanish I and II with approximately 130 to 180 students. The same system did not work at that scale. When six sections submit similar projects within days of one another, the grading mountain grows teeth.
Had I continued, I would have redesigned the model again. Repeated revisions, individualized audio feedback, multiple test versions, and coached reassessment require different systems when seven students participate than when seventy do.
The principles were not wrong. Their implementation depended on class size, course load, planning time, technology, administrative expectations, and the number of separate preparations.
A grading practice is not truly sustainable if it survives only by consuming the teacher.
Multilingualism, Representation, and the Point of It All
Underneath every grading decision was a larger belief: learning multiple languages should not be treated as a luxury.
Parents sometimes joked, “I got an A in Spanish, and all I can say is hola, ¿qué tal?” It was funny because it was familiar, but the joke revealed a serious contradiction. How can a person spend years studying a language, receive a symbol of success, and retain almost nothing they can use?
If students earned excellent grades but could not understand or communicate, the grade had succeeded as a school artifact and failed as evidence of language.
I wanted students to see Spanish as a living language spoken by people of many races, histories, and appearances. This was especially important in a predominantly Black school. Data from SlaveVoyages show that approximately 95 percent of enslaved Africans arriving in the Americas were taken to the Caribbean and South America, while the area that became the United States received only a small share.8
Blackness is not peripheral to Latin American and Caribbean history. Neither are Indigenous identities, mixed-race communities, Asian and Middle Eastern diasporas, migration, conquest, resistance, survival, and cultural exchange.
At the same time, I did not want to replace one simplified story with another. Racism, colorism, and white supremacy did not disappear outside the United States. They developed differently across national and regional contexts.
The goal was to make the picture more truthful. My students deserved to understand that Spanish did not belong to one race, one nation, one accent, or one textbook.
A Democratic and Restorative Gradebook
I wanted my classroom to contain democratic elements. That did not mean every decision was made by popular vote. It meant expectations were visible, systems were explainable, feedback could be discussed, and students had meaningful agency in responding to difficulty.
A student could disagree with a rubric score and produce stronger evidence. A student could recognize a gap and choose to reassess. Teacher authority remained, but it was accountable to stated learning goals.
Restorative practice shaped the same vision. When failure occurred, the question was not only, “What consequence should follow?” It was also, “What needs to be understood, practiced, repaired, or attempted again?”
That question belonged in academic learning as much as it belonged in relationships.
The Best Final Draft I Had
I do not believe this system is a universal template. Some parts would need simplification in larger classes. Some would require stronger department-wide alignment. Some would benefit from better data. Some of my decisions would change based on what I know now.
Equitable grading is not a structure installed once and left untouched. It is an ongoing practice of asking: What is this grade measuring? Does the assessment reflect the learning goal? What barriers are affecting the result? Can students understand how to improve? Is the system transparent? Is it sustainable? Who benefits from its design?
By the end of my high-school teaching years, this was the best version I had created.
It was not perfect. It did not remove every contradiction built into traditional grading. But it allowed me to apply restorative practice, democratic education, cultural responsiveness, universal design, and research in concrete ways.
It created a classroom where rigor did not depend on fear, grace did not require dishonesty, culture was not extra credit, mistakes were not permanent, and a grade could change because a student had changed what they knew and could do.
The ultimate goal remained larger than the number: Could the student understand more Spanish than before? Could they communicate more clearly? Could they interpret another person's meaning? Could they navigate difference with curiosity? Could they see themselves as someone capable of becoming multilingual?
That was the work I was trying to do.
This was my attempt to remake one small part of an education system that was not designed equally for all of the students moving through it.
What have you changed in your grading practices to make learning more accurate, equitable, restorative, and humane? What parts of the system are you still trying to solve?
Research Notes
1. Larry L. Jacoby and Ann Hollingshead found that reading correctly and incorrectly spelled words influenced later spelling accuracy for those words. This supports caution, but the broader literature is mixed enough that the essay treats the practice as a design choice rather than a universal neurological law.
2. ACTFL organizes communication goals around interpretive, interpersonal, and presentational modes and emphasizes real-world communication and cultural understanding.
3. CAST’s Universal Design for Learning Guidelines emphasize predictable learner variability and multiple means of engagement, representation, and action or expression.
4. Thomas R. Guskey and other assessment scholars argue that retakes can support learning when students receive corrective feedback, engage in additional preparation, and complete a genuinely parallel assessment. The practical workload and design conditions remain important.
5. FERPA protects education records maintained by schools. In Owasso Independent School District v. Falvo, the U.S. Supreme Court held that peer-graded work is not an education record before a teacher collects and records the grade.
6. Driessen and colleagues found that students perceived lower anxiety on open-note examinations, but open-note assessment results vary with design, preparation, and context. The essay therefore avoids claiming that open notes affect only students with test anxiety.
7. Guskey argues that percentage scales and averaging create the deeper mathematical problem associated with zeros. A single zero can exert a disproportionate effect because the failing interval occupies far more points than other letter-grade intervals.
8. SlaveVoyages reports that the Caribbean and South America received approximately 95 percent of enslaved Africans arriving in the Americas. Its overview estimates that the United States absorbed approximately 5 percent of arrivals to the Americas.
Bibliography
ACTFL. “NCSSFL-ACTFL Can-Do Statements.” Accessed July 30, 2026. https://www.actfl.org/educator-resources/ncssfl-actfl-can-do-statements.
ACTFL. “World-Readiness Standards for Learning Languages.” Accessed July 30, 2026. https://www.actfl.org/educator-resources/world-readiness-standards-for-learning-languages.
CAST. “CAST Universal Design for Learning Guidelines.” Accessed July 30, 2026. https://udlguidelines.cast.org/.
Driessen, Ellen P., et al. “Evaluating Open-Note Exams: Student Perceptions and Preparation Methods in an Undergraduate Biology Class.” PLOS ONE 17, no. 8 (2022). https://pmc.ncbi.nlm.nih.gov/articles/PMC9387834/.
Guskey, Thomas R. “Giving Retakes Their Best Chance to Improve Learning.” Educational Leadership, April 2023. https://www.ascd.org/el/articles/giving-retakes-their-best-chance-to-improve-learning.
Guskey, Thomas R. “The Case Against Percentage Grades.” Educational Leadership 71, no. 1 (2013): 68-72. https://uknowledge.uky.edu/edp_facpub/19/.
Guskey, Thomas R. “The Unwinnable Battle over Minimum Grades.” Educational Leadership 82, no. 2 (2024). https://www.ascd.org/el/articles/the-unwinnable-battle-over-minimum-grades.
Jacoby, Larry L., and Ann Hollingshead. “Reading Student Essays May Be Hazardous to Your Spelling: Effects of Reading Incorrectly and Correctly Spelled Words.” Canadian Journal of Psychology 44, no. 3 (1990): 345-358.
SlaveVoyages. “A Brief Overview of the Trans-Atlantic Slave Trade.” Accessed July 30, 2026. https://www.slavevoyages.org/blog/-brief-overview-of-the-transatlantic-slave-trade/154.
SlaveVoyages. “Introductory Maps to the Transatlantic Slave Trade.” Accessed July 30, 2026. https://www.slavevoyages.org/blog/introductory-maps-to-the-transatlantic-slave-trade/150.
United States Department of Education, Student Privacy Policy Office. “Family Educational Rights and Privacy Act (FERPA).” Accessed July 30, 2026. https://studentprivacy.ed.gov/ferpa.
United States Supreme Court. Owasso Independent School District No. I-011 v. Falvo, 534 U.S. 426 (2002).

