Consent for the use of cookies and other tools

Tools and cookies used on the website collect information about visitors in anonymized form. Your consents enable us to ensure the functioning of all website features, customize certain content specifically for you, and continuously improve the website by analyzing visits.

Types of Cookies

I consent to the website use of tools, including cookies, which ensure full functionality and an appropriate level of security. I understand that without this, the website cannot offer proper functioning, such as website navigation, customization of appearance, and access to specific parts of the website.

I consent to the website use of tools, including cookies, which collect anonymized data about website visitors. I understand that without this, website administrators cannot analyze site traffic and usage patterns to improve the user experience on the website.

Changes were successfully saved

The news of the past two weeks offers even more versions of such questions and dilemmas. An AI-generated mark differed from the lecturer’s mark by as many as 40 points out of 100. A detector flagged a text as AI-generated but did not explain its decision. An AI system produced initial feedback on a submission containing hidden grading instructions. Once again, the news of the past two weeks shows that checking and explaining AI-generated responses, using them transparently and taking responsibility for them are essential.

More Use of AI Does Not Necessarily Mean More Learning

We begin this issue of the AI-Newsletter with an extensive systematic review and meta-analysis of the use of AI in STEM education. The authors included 85 studies in the systematic review and 49 studies, comprising 59 effects, in the meta-analysis of objectively assessed cognitive outcomes. They concluded that a reliable overall effect of generative AI use on learning could not be confirmed. Larger observed effects were more common in knowledge acquisition than in skills development, and when AI augmented learners’ work rather than replaced their thinking. Users encountered various difficulties when working with AI, including writing effective prompts, cognitive overload, uncritically accepting incorrect information, superficial responses and making poor decisions about which part of a task to delegate to AI. In STEM fields, this was particularly evident in checking equations, writing code and working with probability. The authors conclude that no reliable ways of using AI to improve learning have yet been established. Training in prompt writing, checking AI responses and comparing them with one’s own reasoning is nevertheless valuable, as is additional support for learners with less prior knowledge.

An interesting point in this study is that what matters for the effect is not whether AI is used, but what AI takes over in the learning process and what the student does independently.

This claim is also supported by an article on cognitive resilience - the ability to maintain attention, recall, reasoning and judgement in an environment where AI is constantly available. Among other points, the article argues that universities award qualifications on the basis of demonstrated competence, so assessment must also be based on practical work that requires cognitive friction. Any activity that does not require students to demonstrate acquired knowledge in practice also fails to assess the objectives it claims to assess. Oral examinations, supervised analytical problem-solving and final thesis defences therefore remain valid evidence of learning. This is not simply because it is harder for AI to be present in the classroom in these situations, but because such activities require students to demonstrate abilities that can only be developed through sustained and committed effort. Thinking in real time while receiving feedback thus remains an indication that this effort has actually taken place.

Assessing Student Work

Below, we present several articles that make one point clear: AI can help us prepare assessment questions, create or refine a marking rubric and draft feedback for students, but it should not be entrusted with the final mark, provide explanations on our behalf or assume responsibility for pedagogical decisions.

In recent days, various reports—including Times Higher Education and news from Cardiff University, have highlighted the results of a study in which the authors compared the grades awarded by a lecturer to 50 undergraduate bioscience essays with grades generated by artificial intelligence. AI often graded weaker work too highly and the strongest work too low, while an individual grade differed from the human grade by as many as 40 points out of 100. Because of these differences, the average for the whole group appeared plausible, but the grade awarded to an individual student was unreliable. Yet personalisation should be one of the key advantages of using AI.

Dr Deepshikha at Queen Mary University of London also describes a similar difference between assessment with and without AI. She argues that an AI model can identify superficially fluent writing, structure and the presence of required components, but finds it much harder to assess the development of an argument, the weight of evidence and the significance of an error in a specific learning situation. Drawing on her own experience, she explains that AI tools performed best in assessment, or produced a grade most similar to the human grade, when a well-developed, clear marking rubric was used. She regards the quality of the rubric as the best predictor of agreement between a human and AI. In this case, she also views the AI-generated assessment result as a comment that may help students who rely on rapid feedback. However, she stresses an important distinction: feedback produced by AI tools is merely information or commentary, whereas high-quality feedback should also involve dialogue, be made meaningful and be formulated carefully, thoughtfully and with relationships in mind. At her university, these reflections have led to concrete steps: investing in marking-rubric design, introducing formative assessment and establishing institutional data governance.

security study presented an even more troubling test. By indirectly adding instructions or prompts to a final submission of the kind a student might hand in (the study did not use real student work) the researchers were able to manipulate assessment results. Their aim was to draw attention to the vulnerability of AI-based assessment and the mismatch between security measures and the way AI systems operate. The authors advocate clear guidelines and policies for AI use, improved expertise among educators and others involved in education, and standardised security testing of AI systems’ resistance to such hidden instructions in order to make assessment more robust.

AI Detectors Are Unreliable, but a Return to In-Person Assessment Is Not the Only Answer

When we do not trust AI models to assess student work, we may quickly turn to AI-content detectors. An opinion piece on AI-use policies warns that binary rules based on “permitted use or cheating” repeat the old problems of plagiarism policies. International and multilingual students may be particularly exposed, especially when unusual language use is interpreted as evidence of AI use.

When it comes to scientific images, new AI-detection tools are now considerably better at finding duplications and alterations, but they are not reliable in determining whether an image was artificially generated. The author argues that an image accompanied by a declaration that it was AI-generated may be entirely legitimate, whereas images edited using conventional methods can be misleading. Detectors cannot establish authorship, intent or scientific validity, nor can they determine whether a particular use of artificial intelligence violated the rules.

The authors argue that we are entering a “post-plagiarism” era in which hybrid writing and creation by humans and AI are becoming normal. Instead of assessing a text or image superficially, we should focus on the judgement embodied in it, the process by which it was produced and the responsibility of its authors.

In relation to detection, it is also worth mentioning that the University of Southampton decided to end its contract with a company providing an AI-detection system because it objected to contractual terms concerning the use of student work, although the company disputes this. We cannot draw firm conclusions about the situation from this article alone, but we can ask who stores student submissions in detection tools, for how long, for what purpose and under what conditions. By now, we are probably all aware that contractual fine print has become part of our pedagogical responsibility.

In response to unauthorised AI use, a British think tank issued a call to universities to end unsupervised remote examinations, such as take-home exams, arguing that they affect the public image and credibility of university study. It recommends replacing these examinations with in-person exams and introducing training in the design of modern, AI-resilient assessment tasks. These tasks should be transparent and give everyone involved a clear view of how their work was assessed.

An author addresses a similar theme in an opinion article, arguing that generative AI has exposed certain weaknesses of mass higher education: large lectures, limited dialogue and assessment of the final product rather than the development of thinking. As an alternative, she proposes smaller interdisciplinary groups, peer teaching, discussion, case-based work and the gradual development of texts or other outputs. AI-generated outputs can become objects of analysis, verification and revision. She therefore sees the answer not in more effective surveillance of students, but in making thinking more visible, increasing contact with teachers and peers, and using process-based assessment.

Another article supports redesigning the study process by proposing a five-stage redesign process. First, we review the competencies and assess which parts of the final output can be automated. We then adapt the learning outcomes accordingly, design tasks that reveal the process as well as the final product, set clear boundaries for AI use and, finally, verify whether the student can apply the required knowledge without AI assistance.

An interesting model of ongoing AI-supported student support is presented in an article describing an understanding gap that arises when a teacher infers the whole group’s level of understanding from the responses of only a few students. The author proposes frequent, low-stakes knowledge checks that can be used to offer students personalised learning pathways: additional explanations and examples, short readings, videos and podcasts, or more demanding tasks. AI can help prepare diagnostic questions, repurpose content, create course-specific AI agents and provide rapid, ongoing feedback. All of this can also help involve students who are usually less active.

Write Your Own Draft First, Then Turn to AI for Help

In this section, we examine several concrete examples of the ethical use of AI in the study process.

The first example describes two-stage writing assignments. Students first had to write a draft independently. They then revised it using AI and explained which suggestions they accepted or rejected. The initial draft accounted for 80 per cent of the mark, the AI-assisted revision for 10 per cent and the reflection for the remaining 10 per cent. Students reported that AI brought greater clarity, faster editing and useful language corrections, but they also encountered incorrect subject-specific suggestions and a sense that they were losing their own voice.

In the second example, the author distinguishes between avoiding AI and using AI as a shortcut. Both approaches can leave students without genuine learning. She proposes that AI should act as a discussion partner that gives students basic feedback: for example, identifying a potential strength and weakness in work they are preparing, explaining errors or generating additional practice tasks without actually solving the task facing the student. Here, too, the author encourages students to compare the first and revised versions of their work and reflect on the changes made with AI. Like the AI-Newsletter as a whole, she emphasises cognitive effort, metacognition and critical engagement with AI responses.

In the third example, the authors present an engineering assignment in which students analyse a real process and collect their own evidence of how it works. AI may help explain the instructions, formulate analytical questions, structure the text and improve linguistic clarity. It may not fabricate observations, measurements, sources, screenshots or a process diagram. The assignment requires a visual representation of the process, a check of the data’s internal consistency and a declaration stating which tool the student used, for what purpose and how the results were verified.

The fourth example comes from Slovenia, where a pilot project at DOBA Business School involved 140 students using a general AI tutor with subject-specific knowledge. It could, for example, provide information about exam deadlines, student obligations, reading lists and other technical details, answer organisational questions, and help students read course materials and complete quizzes. In their self-assessments, most students reported being satisfied with the tutor and saw it as a virtual mentor. In addition to the tutor, two further AI personas were introduced. Students used one to check their understanding of the content, while the other used a Socratic approach to help them arrive at answers. In their self-assessments, students reported that using AI saved time and improved their understanding.

As a point of interest, we also include a somewhat different example. The authors describe using a cloned version of a teacher’s voice to produce asynchronous videos, which, according to their account, reduces the time needed to create a recording by as much as 60 per cent. The key workflow consists of a detailed script, an initial recording containing varied speech, audio generation and mandatory human review. One advantage is that individual parts of a lecture can be corrected more easily without recording them again in a studio. The authors place particular emphasis on the speaker’s explicit permission, a purpose defined in advance, appropriate contractual consent and a final review of pronunciation, rhythm and meaning.

From AI Policies to AI Literacy

If we tell students only what is permitted when using AI, we teach them to follow rules rather than to exercise judgement. In an article on AI literacy that goes beyond technical skills, the authors propose four shifts: treating AI as a sociotechnical system; looking beyond purely technical fixes - such as writing effective prompts, identifying hallucinations, recognising bias and evaluating AI-generated content, which are no longer sufficient; developing students’ agency to influence the development of systems and AI-use policies; and considering alternative futures.

UCL and the African Institute for Mathematical Sciences have, to some extent, put this idea into practice by creating a free and adaptable set of teaching materials for learning about responsible AI use, which supplements the Google DeepMind AI Research Foundations curriculum. Across more than 50 hours of activities, the materials connect machine learning with ethics, sociology, philosophy and environmental justice. They provide a useful starting point for anyone who does not want to reduce AI literacy to a list of tools.

So What Should We Do When We Disagree with AI?

The news of the past two weeks makes the answer clear. Human judgement comes first. Not because humans are infallible, but because, at least for now, we are better in considering context, explaining our reasons, listening to students’ responses every day and taking responsibility for our actions, even when those actions may have consequences. AI can offer a warning signal or a second opinion, but it cannot make final decisions for us.

The new University of Ljubljana Guidelines for the Use of Artificial Intelligence were also published today. They establish common principles for using AI in learning, teaching, assessment, research and professional work. Their first principle relates directly to the question raised in this issue of the AI-Newsletter: a human is always responsible for the use of AI and for the final result, while AI does not replace professional, pedagogical or research judgement. The guidelines also address data protection, transparency, academic integrity, equitable access and proportionality of use, and include separate guidance for students, teachers, researchers and professional staff.

An Invitation to Contribute and Continue Learning

As in previous issues of the AI-Newsletter, we invite you to share your experience through a short form if you are experimenting with thoughtful uses of AI in your course, have changed an assignment because of AI or have encountered a dilemma that could also benefit others.

We also invite you to strengthen your AI-literacy competencies through the (currently) four online workshops, which allow you to work independently, at your own pace and at a time that suits you, through topics that have also featured prominently in today’s AI-Newsletter.

Authors: Sanja Jedrinović Čufer, Mateja Bevčič, Eva Kern Nanut, Eva Škraba and Maja Kosmač  
University of Ljubljana Center for the Use of ICT in the Pedagogical Process

Department

UL_Center_mini_logo.png

Center for the use of ICT in pedagogical process (Digital University Center)

Univerza v Ljubljani
Kongresni trg 12
1000 Ljubljana