

In-person tests and exams can play an important role in a course, but they are most effective when used intentionally and in alignment with specific pedagogical goals.
At their best, in-class assessments allow instructors to verify that students can independently perform key intellectual tasks under supervised conditions. This has become particularly relevant in a context where students may rely on a range of external supports, including generative AI tools, when completing work outside of class—an issue explored further in our Creating AI-Resilient Assessments Guide.
However, this does not mean that in-class exams should be the default or primary mode of assessment. Instead, they should be understood as one component within a broader assessment strategy, complementing assignments that allow for deeper, iterative, and collaborative learning. Such an approach is reflected in models like the Two-Lane Approach, which balances in-class demonstration with extended, process-based work.
Situations where in-class assessments are particularly valuable
In-person assessments may be especially useful when you want to:
- Verify independent performance: Ensure that students can demonstrate core skills (e.g., analysis, synthesis, interpretation) without external assistance.
- Assess applied understanding in real time: Evaluate how students interpret, analyze, and respond to materials as they encounter them.
- Complement take-home or project-based work: Balance longer, supported assignments with shorter, controlled demonstrations of learning.
- Reduce reliance on outsourced or AI-generated work: Provide a setting where student thinking can be more directly observed and more authentic work produced.
Common limitations of in-class assessment
Traditional timed, in-person exams can introduce a range of pedagogical and equity challenges. Without careful design, we risk assessing knowledge, skills, and capabilities not directly relevant to learning outcomes. Some common issues include:
Timed exams often reward quick processing and writing rather than sustained, critical thinking. The primary issue here is that what is being measured can also change from deep conceptual understanding to processing speed and surface-level recall when timed tests are the focus (Gernsbacher et al., 2020). That is, when time constraints are set such that students cannot reasonably engage in the analysis, reflection, and revision that deeper understanding requires.
Stress, anxiety, and situational factors (e.g., sleep, illness) can significantly impact results. Zeidner (2007) has provided a comprehensive review of test anxiety in educational contexts and documented that test anxiety in particular is a potent stressor that impairs cognitive performance. Indeed, tests with generous or no time limits can provide a more equitable method of assessment.
For example, as part of her PhD research, Jarvis (1996) compared test scores of college students with and without learning disabilities under standard and extended time conditions. The study found that providing extended time increased scores for students with learning disabilities to levels comparable with nondisabled peers, suggesting that performance under standard time constraints systematically underestimates the true knowledge of some students.
Students working in an additional language or who require more time to articulate ideas may be disadvantaged. For example, Brown (2010) reported that time pressure increases the influence of language and organizational features on essay grades.
Under timed conditions, examiners' judgments are disproportionately shaped by surface-level writing quality (grammar, vocabulary, sentence structure) rather than by the depth or accuracy of the content. This bias disadvantages students with weaker second-language fluency, even when their understanding of the subject matter is strong.
Strict timing and fixed formats can create barriers for students living with disabilities or neurodivergence. As mentioned above, Jarvis (1997) demonstrated that extended time increased test scores for students with learning disabilities to the point of parity with nondisabled peers.
However, Higbee et al. (2008) also discussed universal design principles in higher education and emphasized that accessibility concerns extend beyond time accommodations to include format flexibility, sensory considerations, and multiple means of expression. Fixed, one-size-fits-all exam formats can create systemic barriers for diverse learners.
At UBC, students may can register with the Center for Accessibility to ensure formal accommodations are in place; however, it is still important to ensure our learning design does not rely solely on accommodation processes to make courses accessible.
Wherever possible, accessibility should be built into the structure of the course itself through clear instructions, flexible participation options, accessible materials, multiple ways of engaging with content, and thoughtful assessment design. This helps reduce unnecessary barriers for all students, including those who may not have formal accommodations, may be waiting for documentation, or may not feel comfortable disclosing a disability.
Exams capture performance at one moment, which may not reflect overall learning. More broadly, single-point exams are vulnerable to transient factors that introduce measurement error.
Repeated assessments such as distributed low-stakes quizzes, or optional retakes, can provide a more stable and representative picture of student learning, though equity considerations must guide their design to ensure accessibility and inclusion concerns, like those listed above, are accounted for (Ramming and Mosier, 2018; Supriya et al., 2024).
Timed in-class exams are poorly suited to assessing research, revision, collaboration, and creative work. When these are important learning outcomes, alternative assessment methods, such as portfolios, projects, two-stage exams, or open-book assessments, can help ensure more valid measurements (Supriya et al., 2024; Higbee et al., 2008; Dawson, 2022).
Approaches to designing better in-class assessments
Designing effective in-class assessment requires moving beyond traditional exams toward approaches that both support learning and verify independent performance.
A strong approach combines low-stakes, in-class activities with more structured exams, allowing you to assess student thinking at different points and under different conditions.
Low-stakes, flexible activities
Short, structured activities can be used to assess understanding in real time while reducing pressure and supporting learning. These are especially useful for building skills gradually and giving students opportunities to practice before being formally assessed.
That being said, when designing these kinds of process-based or in-class assessments, instructors should clearly communicate expectations and dates in the syllabus while also considering, from the outset, how students with Centre for Accessibility accommodations may be supported without requiring entirely separate or inequitable assessment pathways.
Design principles
Design activities that target a specific skill or concept in 10–30 minutes, rather than trying to assess everything at once. Research on microlearning and time-restricted activities has found that this can significantly improve knowledge retention compared to longer, traditional sessions by exploiting spacing effects, active retrieval and reduced cognitive load (Aswani, 2025). These bite-sized learning experiences align with working memory constraints and enable distributed practice, which is more effective than massed practice for long-term retention.
Sequentially, ask students to interpret, apply, and evaluate, not simply recall information. Task design can significantly influence the depth of student reasoning.
As Andrew Kreps (2025) found in their examination of a second-semester general chemistry course, cognitively scaffolded activities tend to promote sustained engagement and higher-order reasoning. When tasks target analysis, evaluation and synthesis, students engage more deeply with content and develop transferable thinking skills that can more easily be applied across disciplines and contexts.
Break tasks into steps (e.g., identify > explain > evaluate) to reduce ambiguity and cognitive overload. Structured prompts that break complex tasks into sequential steps, can significantly reduce cognitive load and improve learning outcomes. The key is providing just enough structure to guide thinking without constraining it, using frameworks like backwards design (Morgan and Brooks, 2012) and stepwise scaffolding (Graulich and Caspari, 2020).
Allow brief discussion or pair work, followed by an individual response to verify learning. Peer instruction research provides strong evidence for the effectiveness of brief collaborative interactions. Indeed, the combination of peer discussion followed by instructor explanation can be particularly impactful (Smith et al., 2011).
Notably, strong performers may not benefit to the same degree from instructor-only approaches, highlighting the importance of peer discussion even for top students. That being said, rather than framing two-stage collaborative assessments as a way to simply improve performance, it may be more useful to emphasize how they preserve individual accountability while adding a structured peer-learning component by asking students to first articulate their own reasoning and then compare, defend, and refine that reasoning with others.
Each activity should target one clearly defined skill (e.g., argument evaluation, concept application), rather than trying to assess multiple outcomes at once.
The literature clearly shows that activities aligned with a single, clearly defined learning objective are more effective than those attempting to address multiple outcomes simultaneously (McCann, 2017). As a pedagogical strategy, constructive alignment ensures that what students do (activities), what they are asked to demonstrate (assessment) and what they are intended to learn (outcomes) are in harmony, reducing confusion and cognitive load (Rifan and Latif, 2025).
Provide 2–3 clear indicators of what a strong response includes (e.g., “uses course concept accurately,” “provides specific evidence”).
After all, research on rubrics and formative assessment clearly demonstrate that clear learning targets and compact criteria support student learning and self-assessment. Simple rubrics can be used for lesson-sized formative assessment to develop longer-term goals, providing practical guidance for students (Andrade and Brookhart, 2026). While comprehensive rubrics might have their place, concise criteria are particularly effective for short, focussed activities because they reduce cognitive load and allow students to better concentrate.
Use formats that allow you (or peers) to quickly assess responses without heavy grading. For example, ask for short written responses, checklists, or structured outlines that can be reviewed in minutes. Whether this feedback is provided by an instructor or peer, the key is to allow for multiple practice-feedback cycles within a limited timeframe, which has been shown to significantly improve recall of course material, enhance metacognition, increase self-reported mastery, boost attentional control, and reduce anxiety (Higham et al., 2022).
Develop a small set of repeatable activity types (e.g., critique, apply, outline) that students become familiar with over time. Such familiarity reduces extraneous cognitive load, can enable procedural fluency, and allow students to focus on content and higher-order thinking rather than decoding new instructions (Aswani, 2025; Wagner-Loera, 2018).
Sequence activities so they increase in complexity or mirror components of larger assignments. Scaffolding should be gradually reduced and eventually removed as students gain understanding, a process known as fading.
This gradual release of responsibility ensures that students develop independence while maintaining support during initial learning phases. Sequencing from scaffolded, supported tasks to less-supported, more complex tasks produced higher performance and supports skill development over time (Tucker et al., 2021; Morgan and Brooks, 2012).
Examples
Students read a short article (e.g., opinion piece or news article) and identify the central claim, evaluate the strength of evidence provided, and suggest one specific improvement.
Example (English): Students analyze a short critical essay or review, identifying the author’s central interpretation of a text and evaluate how effectively textual evidence is used to support that interpretation, before suggesting a more compelling reading or additional evidence.
Students apply a course concept to a real-world example.
Example (Geography): Students apply gentrification theory to analyze a short news article on housing in Vancouver, using the concept to interpret patterns of change and displacement.
Students analyze a short text, dataset, map, or image and explain its significance using course frameworks.
Example (Sociology): Students analyze a short dataset or media excerpt and use a theoretical framework (e.g., social stratification or intersectionality) to explain its broader social significance.
Students are given a prompt and asked to produce a clear argument outline (thesis + key points + evidence).
Example (History): Students are given a primary-source based prompt and asked to outline an argument (thesis, key points, and supporting evidence from the text) that situates the source within its broader historical context.
Students revise a paragraph to improve clarity, argumentation, or use of evidence.
Example (Sociology): Students revise a short argumentative paragraph to improve logical coherence, clarify the thesis, and strengthen its use of evidence or reasoning.
Students identify the strongest counter-argument to a claim and briefly respond to it.
Example (Political Science): Students identify the strongest counter-argument to a policy claim (e.g., on climate or immigration) and briefly respond using relevant theoretical or empirical evidence.
In-class exams
When using formal in-class exams, the goal is not simply to test recall under time pressure, but to assess applied understanding and independent thinking in a controlled environment.
Design principles
Use questions that require students to analyze, synthesize, or evaluate instead just reproducing information. Jensen et al. (2014) found that students tested with high-level questions throughout the semester demonstrated significantly higher performance on both low-level and high-level final exam items compared to students tested primarily on recall.
Design exams so a well-prepared student could reasonably complete them in about 50% of the allotted time, allowing space for planning and revision. Worrisomely, research indicates that when time pressure is not accounted for, it can raise concerns surrounding equity.
In a controlled field experiment at an Italian university, De Paula et al. (2016) found that time pressure caused statistically significant drops on both verbal and numerical tasks; although, this was driven almost entirely by female students in the study, with no significant effect on males. What is more problematic, though, is that this effect was primarily driven by a strong negative impact on female students’ performance, while male students showed no significant effect.
Allow students to choose between questions or cases to reduce the impact of isolated gaps in knowledge. Even though research on student choice in assessment contexts shows mixed results regarding performance, there is consistent evidence that it positively impacts engagement and comfort.
For example, Campbell and Donahue’s (1997) analysis found that students did not improve performance, but that they perceived the assessment as easier, suggesting psychological benefit without performance gains. Similarly, Hott et al. (2023) did not find improvements to performance in a computer science course but did find that choice reduced testing stressors.
Indicate how much time to spend on each question and what a strong response should include. For example, Hott and Pettit (2023) found that timers and clear time-allocation cues help students manage pacing, thought visible timers can also raise anxiety. Offering timer choices may be one option, as this was shown to improve comfort.
Introduce preparation, revision, or multiple sittings to better reflect how learning actually occurs. Two-stage examinations have been extensively studied across multiple disciplines and generally show positive results (Levy et al., 2019; Callaghan et al., 2025).
If students will be using their own devices, require the use of LockDown Browser to restrict access to prohibited websites and supports (e.g., generative AI tools).
Instructors may wish to book one of our computer labs to administer digital assessments in a controlled environment. Arts ISIT is also currently piloting a new model for asynchronous, invigilated exams with ORCA (previously the Computer Based Testing Facility). If you are interested in piloting this, please contact Arts ISIT for more information.
Examples
Students work with short materials and produce a structured response.
Example: Read 1–2 short texts, identify key arguments or themes, and compare perspectives and apply course concepts.
Students apply course concepts to a real-world situation.
Example: Students choose from a few news articles or case studies and analyze one using course frameworks while evaluating implications or outcomes.
The exam is divided into several short, targeted tasks.
Example: Short analysis (15 mins.); concept application (15 mins.), argument critique or revision (15-20 mins.).
Students complete the exam in two parts with an opportunity to revise.
Example: In session one they write initial response under time constraints, and in session 2 they revise, extend, or respond to feedback.
Questions are released 24-48 hours, students are asked to take notes or outlines, and then each must write a response in class without external support.
In-class assessment can play a valuable role in your course, but its effectiveness depends on how intentionally it is designed. By combining low-stakes activities with thoughtfully structured exams, you can both support student learning and verify independent performance in meaningful ways.
Rather than defaulting to traditional formats, consider what you truly want students to demonstrate, and then choose approaches that align with those goals. Even small changes in how in-class assessments are designed can lead to more equitable, authentic, and effective evaluation of student learning.
References
Andrade, H.L., & Brookhart, S.M. (2026). Using Rubrics for Teaching and Learning (1st ed.). Routledge. https://doi.org/10.4324/9781003582649
Aswani T D. (2025). Microlearning Moments: The Science of Knowledge Retention in Bite-Sized Educational Experiences. International Journal of Education and Pedagogy (IJEP), 1(4), 116–121. https://doi.org/10.5281/zenodo.17318186
Brown, G. T. L. (2010). The validity of examination essays in higher education: Issues and responses. Higher Education Quarterly, 64(3), 276–291. https://doi.org/10.1111/j.1468-2273.2010.00460.x
Callaghan, E. S., Hawkins, L. K., & Colvard, M. J. (2025). Active learning through flexible collaborative exams: Improving assessments across disciplines. Active Learning in Higher Education. Advance online publication. https://doi.org/10.1177/14697874251344293
Campbell, J. R., & Donahue, P.L. (1997). Students selecting stories: the effects of choice in reading assessment : results from the NAEP reader special study of the 1994 National Assessment of Educational Progress. Washington, D.C. (555 New Jersey Ave., NW, Washington, D.C. 20208-5574): U.S. Dept. of Education, Office of Educational Research and Improvement, National Center for Education Statistics.
Dawson, P. (2022). Inclusion, cheating, and academic integrity. In T. Bretag (Ed.), Academic integrity in the age of online learning (pp. 181–194). Routledge. https://doi.org/10.4324/9781003293101-13
De Paola, M., and F. Gioia. 2016. “Who Performs Better Under Time Pressure? Results From a Field Experiment.” Journal of Economic Psychology 53:37–53. https://doi.org/10.1016/j.joep.2015.12.002
Gernsbacher, M. A., Soicher, S. L., & Becker-Blease, A. J. (2020). Four empirically based reasons not to administer time-limited tests. Translational Issues in Psychological Science, 6(2), 175–190. https://doi.org/10.1037/tps0000232
Graulich, N., & Caspari, I. (2021). Designing a scaffold for mechanistic reasoning in organic chemistry. Chemistry Teacher International, 3(1), 19–30. https://doi.org/10.1515/cti-2020-0001
Higbee, J. L., Ed, Goff, E., Ed, & Minnesota Univ., Minneapolis. Center for Research on Developmental Education and Urban Literacy. (2008). Pedagogy and student services for institutional transformation: Implementing universal design in higher education.
Higham, P. A., Zengel, B., Bartlett, L. K., & Hadwin, J. A. (2022). The benefits of successive relearning on multiple learning outcomes. Journal of Educational Psychology, 114(5), 928–944. https://doi.org/10.1037/edu0000693
Hott, B. L., Sherriff, M., & Pettit, R. (2023). Providing a choice of time trackers on online assessments. In Proceedings of the 54th ACM Technical Symposium on Computer Science Education (SIGCSE 2023) (pp. 1106–1112). Association for Computing Machinery. https://doi.org/10.1145/3545945.3569776
Ives, J. (2014, July 30-31). Measuring the Learning from Two-Stage Collaborative Group Exams. Paper presented at Physics Education Research Conference 2014, Minneapolis, MN. Retrieved April 1, 2026, from https://www.compadre.org/Repository/document/ServeFile.cfm?ID=13464&DocID=4063
Jarvis, K. A. (1996). Leveling the playing field: A comparison of scores of college students with and without learning disabilities on classroom tests
Jensen, J. L., McDaniel, M. A., Woodard, S. M., & Kummer, T. A. (2014). Teaching to the test…or testing to teach: Exams requiring higher order thinking skills encourage greater conceptual understanding. Educational Psychology Review, 26(2), 307–329. https://doi.org/10.1007/s10648-013-9248-9
Kreps, A. J. (2025). From materials to meaning: Exploring the impact of cognitive scaffolding on group dynamics and sensemaking (Doctoral dissertation, University of Iowa). https://doi.org/10.17077/etd.006755
Levy, B. L., Rubel, M. J., & Lester, J. (2019). Two-stage examinations: Can examinations be more formative experiences? Active Learning in Higher Education, 20(3), 209–221. https://doi.org/10.1177/1469787418801668
McCann, M. (2017). Constructive alignment in economics teaching: A reflection on effective implementation. Teaching in Higher Education, 22(3), 336–348. https://doi.org/10.1080/13562517.2016.1248387
Morgan, K., & Brooks, D. W. (2012). Investigating a method of scaffolding student-designed experiments. Journal of Science Education and Technology, 21(4), 513–522. https://doi.org/10.1007/s10956-011-9343-y
Ramming, C. H., & Mosier, R. (2018, June), Time Limited Exams: Student Perceptions and Comparison of Their Grades versus Time in Engineering Mechanics: Statics Paper presented at 2018 ASEE Annual Conference & Exposition, Salt Lake City, Utah. 10.18260/1-2—31144
Rifan, N. M., & Latif, A. A. (2025). Constructive alignment as a framework for enhancing motivation and higher-order thinking in science classrooms: A systematic synthesis. International Journal of Research and Innovation in Social Science, 9(3), 8078–8086. https://doi.org/10.47772/ijriss.2025.903sedu0605
Smith, M. K., Wood, W. B., Krauter, K., & Knight, J. K. (2011). Combining peer discussion with instructor explanation increases student learning from in-class concept questions. CBE—Life Sciences Education, 10(1), 55–63. https://doi.org/10.1187/CBE.10-08-0101
Supriya, K., Ballen, C. J., Jeong, J. L., Guo, D., & Brownell, S. E. (2024). Optional exam retakes reduce anxiety but may exacerbate score disparities between students with different social identities. CBE—Life Sciences Education, 23(1), ar6. https://doi.org/10.1187/cbe.21-11-0320
Tucker, T., & Mercier, E., & Shehab, S. (2021, July), The Impact of Scaffolding Prompts on Students’ Cognitive Interactions During Collaborative Problem Solving of Ill-structured Engineering Tasks Paper presented at 2021 ASEE Virtual Annual Conference Content Access, Virtual Conference. 10.18260/1-2-37871
Wagner-Loera, D. (2018). Flipping the ESL/EFL classroom to reduce cognitive load: A new way of organizing your classroom. In J. Mehring & A. Leis (Eds.), Innovations in flipping the language classroom: Theories and practices (pp. 159–174). Springer. https://doi.org/10.1007/978-981-10-6968-0_12
Zeidner, M. (2007). Test anxiety in educational contexts: Concepts, findings, and future directions. In P. A. Schutz & R. Pekrun (Eds.), Emotion in education (pp. 165–184). Academic Press. https://doi.org/10.1016/B978-012372545-5/50011-3


