Dr Olajide JolugboLearning Ecosystem Architect
Back to Insights
AI in Education8 min read17 July 2026

AI in Education and Assessment

A practical look at where AI genuinely strengthens teaching and assessment, and where it introduces risks to integrity and trust

The educational case for artificial intelligence is frequently overstated because efficiency is mistaken for learning. Faster lesson planning, instant feedback, polished submissions and automated scoring may reduce effort, but they do not necessarily improve understanding, judgement or competence.

The critical question is not whether AI can perform an educational task. It is whether delegating that task strengthens learning, weakens essential intellectual effort or compromises the validity of the decisions made from it.

AI creates educational value when it extends the capacity of teachers and learners without concealing who performed the thinking. It becomes dangerous when fluent output is accepted as evidence of knowledge, automated judgement is treated as objective or institutional enthusiasm moves faster than educational scrutiny.

Efficiency is not educational value

AI can reduce the time required to prepare lesson outlines, generate examples, adapt reading materials and produce preliminary feedback. Early adopter schools and colleges interviewed for the Department for Education particularly identified workload benefits in lesson planning, resource creation and administrative work. These are legitimate gains, but they remain gains in production.

A lesson plan generated in seconds is not necessarily a well designed lesson. Automated feedback is not necessarily accurate, developmental or appropriate to the learner. A resource differentiated by reading level may simplify vocabulary while damaging conceptual precision.

The danger is that AI makes weak material inexpensive to produce and easy to distribute. Without subject expertise and quality assurance, efficiency simply increases the speed at which error, superficiality and generic teaching reach learners.

AI should therefore be used most confidently where the output remains provisional and the professional retains control. It can challenge a teacher’s first idea, generate alternative scenarios or reduce repetitive preparation. It should not determine what is educationally appropriate merely because it produces plausible material quickly.

AI works best when constrained by pedagogy

Evidence that AI can strengthen learning exists, but it does not support unrestricted use of general purpose chatbots.

A randomised controlled trial in an undergraduate Harvard physics course found that students using a specially developed AI tutor achieved greater learning gains in less time than students receiving an active classroom lesson. However, the tutor was not simply an open chatbot. It was carefully structured around active learning, cognitive load, sequential scaffolding, verified solutions and targeted feedback. The researchers deliberately limited reliance on the language model because of the risk of inaccurate outputs.

The result is important, but its interpretation should be disciplined. It demonstrates the potential of a deliberately engineered instructional system in a defined subject and short intervention. It does not establish that giving students unrestricted access to a commercial chatbot will produce comparable learning.

Tutor CoPilot provides an equally revealing example. In a study involving more than 700 tutors and 1,000 learners, tutors received real time AI suggestions for responding to mathematical errors. Students taught by tutors using the system were four percentage points more likely to master topics, with larger gains among learners taught by less highly rated tutors.

Its strength lies in supporting human expertise rather than removing it. Tutors retained the contextual knowledge and authority to decide whether a suggestion was appropriate. The AI expanded access to stronger pedagogical prompts, but a human remained accountable for the interaction.

Pedagogically constrained AI may improve particular learning processes. Unstructured access may produce very different outcomes.

Better performance with AI may conceal weaker learning

The most serious error is to confuse successful task completion with the development of independent capability.

A large field experiment in secondary mathematics found that access to GPT based assistance improved students’ performance during practice. Yet learners using an unrestricted version subsequently performed worse when the assistance was removed. A safeguarded tutoring version reduced this harm by directing students through the reasoning rather than simply supplying answers.

This distinction has major implications for both teaching and assessment. When AI drafts an essay, solves a problem, writes code or structures an argument, the resulting product may be stronger while the learner’s capability remains unchanged. The output demonstrates what the learner and the system produced together. It does not automatically demonstrate what the learner can understand, defend or reproduce independently.

AI can therefore create an illusion of improvement. Learners submit more polished work, tutors encounter fewer basic errors and completion becomes faster. Yet the cognitive work that produces durable learning may have been displaced.

Teaching should consequently distinguish between AI as scaffolding and AI as substitution. The difference is not determined by whether AI was used. It is determined by what thinking remained with the learner.

AI as scaffolding

Prompts explanation, questioning, comparison and revision. The learner remains responsible for the thinking.

AI as substitution

Performs the intellectual work while the learner approves or lightly edits the result. Capability is not exercised.

Assessment integrity is a design problem

Generative AI has not merely introduced a new form of cheating. It has exposed the weakness of assessment tasks that treat an unsupervised written product as unquestionable evidence of individual competence.

Ofqual warns that AI generated text, images, audio, video and code may not reflect a learner’s own knowledge or understanding. This threatens assessment validity and the reliability of results, particularly where acceptable use is unclear or difficult to monitor. Ofqual also cautions that technological detection should be treated only as one source of evidence because false positive and false negative results remain possible.

An arms race between generation and detection is therefore an inadequate strategy. It risks accusing learners on uncertain evidence while leaving fundamentally vulnerable assessments unchanged.

A stronger response is to reconsider what evidence of capability should look like. This may include supervised demonstrations, oral questioning, staged submissions, authentic data, reflective justification, practical observation and comparison between AI supported and independent performance.

The University of Sydney’s two lane approach illustrates this direction. Its model combines secure assessments, in which learners demonstrate course outcomes without external assistance, with open assessments in which appropriate AI use is permitted and acknowledged.

The approach is more credible than either prohibition or unrestricted acceptance because it recognises two legitimate requirements: graduates must be capable of working with contemporary tools, but qualifications must still certify capabilities that belong to the graduate.

Nevertheless, two lanes do not automatically guarantee validity. Secure assessment can become an overreliance on traditional examinations, while open assessment can still reward access to superior paid tools, stronger prompting skills or undisclosed external support. Each task must remain aligned with the capability it claims to assess.

Automated marking shifts rather than removes judgement

AI assisted marking is often presented as a solution to workload and feedback delays. The risk is that consistency becomes confused with fairness.

A system can apply the same rule repeatedly and still apply a poor rule. It may favour conventional language, penalise unusual but valid reasoning or overlook the disciplinary meaning that an experienced assessor would recognise. Bias does not disappear because a decision is automated. It becomes embedded within the model, rubric, training data and institutional workflow.

AI may be appropriate for preliminary categorisation, low stakes practice, identifying missing elements or suggesting feedback for professional review. It is much harder to justify as the final authority over high stakes grades, progression or certification.

Assessment decisions must remain explainable and contestable. A learner should not be disadvantaged because an institution cannot explain how an automated judgement was produced or because staff assume that a numerical output is more objective than professional expertise.

Corporate L&D faces the same validity problem

The risks are not confined to schools and universities. Corporate Learning and Development teams are increasingly using AI coaches, simulations and skills assessments.

Bank of America reports that its Academy uses AI conversation simulators to allow employees to practise client interactions and receive immediate feedback. Employees completed more than one million simulations during 2024.

This demonstrates scale and access to repeated practice. It does not, by itself, establish that employees developed better judgement or improved real client outcomes.

An AI simulator may provide psychologically safe rehearsal and consistent opportunities to practise. However, its assessment is only as credible as the behaviours it has been designed to recognise. Employees may learn to satisfy the simulator rather than handle the ambiguity, emotion and unpredictability of an authentic conversation.

Corporate L&D must therefore resist using simulation scores as direct evidence of workplace competence. The stronger evaluation is whether employees apply the capability appropriately in practice, whether managers observe improved behaviour and whether service or performance outcomes change. One million completed simulations is a measure of activity. It is not automatically a measure of learning transfer.

Trust cannot be added after implementation

AI use affects trust when learners do not understand what is permitted, how their data is used or whether their work will be judged by a machine.

Jisc’s 2025 research found that students continued to report uncertainty about acceptable AI use despite many institutions having issued guidance. Students also raised concerns about unequal access to paid tools, misinformation, privacy, skills loss and inconsistent expectations between courses and tutors.

This exposes a governance failure. Publishing a general policy is insufficient when assessment expectations remain ambiguous at task level. Trust requires explicit decisions about:

  • Which uses are permitted for each assessment.
  • What learners must disclose.
  • Which capabilities must be demonstrated independently.
  • Whether AI is involved in marking or monitoring.
  • How automated outputs are reviewed and challenged.
  • What personal or intellectual property data may enter external systems.

UNESCO similarly argues for a human centred approach in which AI use protects human agency, inclusion and public accountability rather than allowing commercial technology to determine educational practice.

The test is what remains human

The strongest uses of AI preserve the difficult work through which people learn. They create additional practice, offer timely prompts, help teachers examine alternatives and support feedback that a qualified professional can verify.

The weakest uses remove the very thinking an assessment is intended to reveal. They generate finished answers, automate high stakes judgement, conceal uncertainty and convert learner behaviour into data without meaningful consent or challenge.

Institutions and corporate organisations should therefore ask:

What capability is AI strengthening, what human effort is it replacing, and what evidence shows that learning still belongs to the learner?

Without defensible answers, AI may increase productivity while weakening education. It may make assessment faster while making qualifications less trustworthy.

The goal should not be to introduce as much AI as possible. It should be to use AI only where it strengthens human teaching, protects valid assessment and leaves learners more capable of thinking and acting without it.

References

  1. Attewell, S. (2025). Student perceptions of AI 2025. Jisc.
  2. Bank of America. (2025, April 8). AI adoption by BofA’s global workforce improves productivity, client service.
  3. Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences, 122(26), e2422633122. https://doi.org/10.1073/pnas.2422633122
  4. Department for Education. (2025). Generative artificial intelligence in education. UK Government.
  5. Kestin, G., Miller, K., Klales, A., Milbourne, T., & Ponti, G. (2025). AI tutoring outperforms in class active learning: A randomised controlled trial introducing a research based design in an authentic educational setting. Scientific Reports, 15, 17458. https://doi.org/10.1038/s41598-025-97652-6
  6. Miao, F., & Holmes, W. (2023). Guidance for generative AI in education and research. UNESCO.
  7. Ofqual. (2026). Artificial intelligence malpractice and assessment: Advice note. UK Government.
  8. University of Sydney. (2024, November 27). University of Sydney’s AI assessment policy: Protecting integrity and empowering students.
  9. Wang, R. E., Ribeiro, A. T., Robinson, C. D., Loeb, S., & Demszky, D. (2025). Tutor CoPilot: A human AI approach for scaling real time expertise (EdWorkingPaper No. 24-1054). Annenberg Institute at Brown University. https://doi.org/10.26300/81nh-8262
Have a topic to discuss

Interested in a defensible AI and assessment strategy?

I help institutions and corporate L&D teams design AI use that strengthens teaching, protects assessment validity and keeps learning accountable to the learner.

Discuss a Project