Increasingly Capable AI, Increasingly Demanding Evaluation
ChatGPT
Date of publication:
AI is becoming increasingly capable in learning, research, employment, and the independent execution of tasks, which brings to the forefront the question of how to evaluate its results and maintain human accountability. This newsletter presents new approaches to evaluation and AI literacy, the use of AI as a research collaborator, and its impact on the labor market. It also raises questions about identifying content created by AI and the risks posed by increasingly autonomous systems. Finally, it highlights a somewhat different use of AI—in the reinterpretation of works of art.
Assessment during AI: from detection to proving understanding
The article Students ‘Confused About the Purpose of Assessment’ in the Age of AI presents interviews with early-career academics at British universities. They observe that generative AI is exacerbating existing challenges in academic writing. Students are increasingly viewing written assignments primarily as a product to be submitted, rather than as a process through which they develop critical thinking, analysis, synthesis, and argumentation. The authors emphasize that it is therefore necessary to explain more clearly why writing is, in and of itself, part of the learning process.
At the same time, an increasing number of U.S. universities are moving away from plagiarism detectors as the sole evidence of unauthorized use of AI. Yale, Vanderbilt, Johns Hopkins, and Indiana University have restricted or discouraged their use, while several institutions have disabled Turnitin’s AI detection feature. The reasons include inconsistent and false-positive results, as well as poorer performance when analyzing texts written by students for whom English is not their first language. Instead of searching for cheaters, attention is therefore shifting even more toward assignments in which students must demonstrate their work process, thinking, and understanding.
The problem becomes even more challenging in online learning. In an article titled A.I. Agents Are Taking Entire Online Courses for Cheating Students, The New York Times describes agents that can perform activities in online courses on a student’s behalf, take tests, write assignments, and participate in online discussions. Such use goes beyond assistance with individual assignments and allows a large part of the learning process to be transferred to the AI. While online learning significantly increases access to education, its asynchronous and digital nature makes it difficult to verify who actually completed the work.
One possibility is to incorporate more dialogue into the assessment process. The authors of the commentary From detection to dialogue: Assessment with AI voicebots? propose voice-based AI bots that would encourage students, through questions and follow-up questions, to explain their decisions and reasoning. Such an approach could also make oral assessments more accessible to larger groups, particularly in the context of formative assessment. However, they emphasize that quality does not depend on the tool itself, but rather on the design of the task; at the same time, equal access, data protection, and human oversight are also important.
The potential of such assessment formats is also explored in the study Asynchronous oral assessments: enhancing integrity, engagement, and communication in the AI era. Students received questions and recorded short oral responses within a limited time frame. In accounting, the correlation between performance on the oral exam and subsequent written exam was not statistically significant; however, in data analytics, students achieved higher scores on the oral assessment and reported more active preparation methods. The authors view asynchronous oral assessment primarily as a complement to other forms of assessment.
How to Incorporate AI into Learning Without It Taking Over the Student's Thinking
The study Designing Epistemic Partnerships with Generative AI followed 31 first-year students over a two-year period in a course on scientific argumentation. When activities were designed to require students to structure arguments, debate with peers, and critically evaluate the AI’s responses and limitations, the proportion of more challenging interactions with ChatGPT increased from 12.5% to 47%. Students thus shifted from using this AI tool as a source of answers to using it as a conversation partner in constructing and evaluating knowledge. The authors emphasize that AI does not, by itself, become a high-quality learning partner. The key lies in how the learning activity is designed.
AI literacy, therefore, is not merely knowledge of the rules of permitted use. The author of the article AI literacy cannot be bolted on to first-year courses incorporated it directly into two first-year courses in teacher education. Students can use AI to navigate a complex text, explain a concept, or find an example, but they then return to the original source, verify the answer, and assess what has been simplified or omitted. The responsibility for understanding and professional judgment thus remains with the student.
Just how challenging this judgment is in practice is demonstrated by a study of AI literacy among 161 prospective teachers in China. They rated their general knowledge of AI relatively highly, but the results were inconsistent across individual dimensions and subject areas. Interviews with nine students revealed, above all, a gap between their understanding of general ethical issues—such as bias, privacy, and over-reliance—and their ability to apply these principles in specific teaching situations.
In online learning, AI can also provide practice opportunities that are otherwise difficult to organize. The author of the article Using AI to create the practice opportunities online students need developed an AI conversation partner to simulate challenging conversations and another tool for communication feedback. Students thus practiced listening, responding to opposing viewpoints, and handling conflict situations. The purpose of the AI here is not to generate a response, but to provide a safe space for repeated practice of skills that are developed primarily through hands-on experience.
AI as a research associate
The use of large language models (LLMs) is also rapidly expanding in the preparation of research grant proposals. An analysis of 5,700 confidential proposals and 131,000 funded projects from the U.S. agencies NSF and NIH showed that proposals with a higher perceived use of LLMs were less distinct in content from recently funded projects. At the NIH, the use of LLMs was associated with a higher probability of funding and more early-stage publications, though these were generally less impactful; at the NSF, these associations were not observed. The authors therefore warn of the potential for a reduction in the diversity of research ideas.
Agent-based AI is also taking over part of the technical work that underpins modern research. OpenAI’s report on scientific computing presents eight projects, mostly in the life sciences, in which coding agents helped update, optimize, and rewrite research software. As a result, researchers spent less time on implementation and more on setting goals and verifying results. At the same time, the report emphasizes that agents are not yet able to reliably assess the scientific validity of their results, so human verification remains essential.
An interesting example of human-AI collaboration also comes from cryptography. The Claude model helped uncover a vulnerability in the experimental HAWK digital signature system, which was designed to remain secure even against future quantum computers. Researchers were already very close to a successful attack, and AI helped identify the missing step; a researcher independently achieved a similar result that same week. Due to the identified vulnerability, HAWK was withdrawn from the standardization process; this case demonstrates that AI can make a significant contribution to research, while the search for the right research direction and expert judgment remain in the hands of researchers.
Researchers went a step further by using the genomic language models Evo1 and Evo2 to design functional bacteriophages—viruses that attack bacteria. Out of several thousand proposed genomes, they synthesized and tested nearly 300 in the laboratory; 16 of these produced functional bacteriophages, and their combination overcame the resistance of two strains of E. coli. The researchers deliberately excluded viruses that can infect humans, animals, or plants from the training data; nevertheless, this achievement raises important questions regarding biosafety and the regulation of generative genome design.
When a candidate uses AI and when AI evaluates a candidate
AI is also rapidly changing how graduates enter the labour market. Research presented in the article AI reshapes record race for graduate jobs in 2026 – Study shows that 88% of surveyed candidates in South Africa used AI in their job search, primarily to prepare for interviews and to write resumes and motivation letters. At the same time, 37% of employers already use AI in the candidate selection process. This reveals an interesting discrepancy: while candidates often use AI themselves, they are much more hesitant about employers using it to evaluate them.
Predictions of mass job losses due to AI have not materialized so far. The article AI was supposed to destroy jobs. Where’s the carnage? notes that the biggest changes are currently taking place within jobs themselves: individual tasks are being automated, employers are reducing staffing for certain tasks, and they are raising their expectations regarding AI skills. According to one of the studies presented, 74% of employers already consider AI skills to be an important advantage or requirement, which is why the ability to work with AI is becoming increasingly important when entering the labor market.
How will we know what is real and what was created by AI?
As of August 2, new transparency requirements under the Artificial Intelligence Act are in effect in the EU. For certain systems, users must be informed that they are interacting with AI; deepfakes must be labeled; and content created or modified by AI must also include machine-readable tags that enable its identification. The purpose of these rules is to reduce the potential for deception and make it easier to distinguish between authentic and synthetic content.
Anthropic explains how such labeling might work for text in the article How Claude’s text watermark works. In future versions of Claude, the model will not add hidden characters to the text; instead, the choice of words will create a pattern imperceptible to humans that can be verified using a special key. Anthropic is also developing an interface for detecting these watermarks. However, such a mark does not prove that the entire text was written by AI—it can only indicate the likelihood that Claude was involved in its creation or extensive editing.
The issue of identifying synthetic content can have very concrete consequences. In the United Kingdom, the Report Remove service received 420 reports from children regarding fake or manipulated explicit images in the first half of the year alone—more than in all of 2025. AI and digital clothing-removal tools make it possible to create such images even from completely ordinary photographs. This example shows that deepfakes are not just a matter of misinformation, but also of security, privacy, and new opportunities for abuse and extortion.
When AI doesn't just respond, but starts acting on its own
Security testing by the British AI Security Institute demonstrates how greater autonomy can lead to unexpected behavior. During a cyber test, agents powered by models from Anthropic and OpenAI performed actions that were not explicitly instructed while attempting to achieve their stated goal. One attempted to inject malicious code into an open-source project, then created fake online identities and targeted real developers to persuade them to accept it. Although no damage was caused, the incident highlights a new type of risk: a more autonomous system may, in pursuit of its goal, independently choose an unforeseen and deceptive path.
When AI “illuminates” art with a solar eclipse
To mark the total solar eclipse we were able to observe last week, Madrid’s Prado Museum created a digital project titled El Prado a la luz del eclipse, in which it reimagined some of its masterpieces using artificial intelligence and traditional graphic animation. The one-minute video shows how the changing light, shadows, and colors during the eclipse would transform the scenes in the paintings. This example demonstrates a somewhat different use of AI: as a tool for reinterpreting and offering a new perspective on existing works of art.
Invitation to educators
If you have explored an interesting use of AI in education at your faculty, in a course, or as part of your professional work, we cordially invite you to share your example with us via the form for collecting examples of best practices. We look forward to reviewing the submissions and including selected contributions in future issues of the AI Newsletter so that, together, we can build an overview of useful and thoughtful approaches in the university setting.
Authors: Maja Kosmač, Sanja Jedrinović Čufer, Mateja Bevčič, Eva Kern Nanut, Eva Škraba, University of Ljubljana Center for the use of ICT in the pedagogical process
Department
Center for the use of ICT in pedagogical process (Digital University Center)
Univerza v Ljubljani
Kongresni trg 12
1000 Ljubljana