
cApStAn at the ITC Conference in Auckland
International Test Commission (ITC) Conference, Auckland, New Zealand, June 30-July 3, 2026
More research data or more responses to ethical concerns?
Conference Website: https://www.itc2026auckland.com
One of the rewarding aspects of cross-pollination between researchers and field practitioners is that one gets to attend the biennial International Test Commission (ITC) Conference. The 2026 edition took place in Auckland, New Zealand, from June 30 to July 3. That’s about as far from Brussels as it gets: if you started drilling a hole in Brussels and drilled on right through the centre of the Earth, you’d eventually come out in New Zealand. For a Belgian, the word “antipodes” acquires its full meaning in Aotearoa. I was lucky enough to be the one representing cApStAn at the ITC Conference this year, and it has been a particularly inspiring week, both intellectually and philosophically.
Between the sessions, keynotes and symposia, one colleague (a senior researcher who once introduced himself as follows: “my name is […] and I tend to analyse data”.) argued that, considering that the ITC event is a research conference, there was a dearth of data and an abundance of roadmaps, ethical discussions and practical recommendations. He advocated for more research data. Another colleague, who leads an institute that manages admission exams to higher education, had the opposite view: “I can find the data in the papers and publications, I am interested in stories, ideas, experiences and in the human side of all things assessment. I don’t come all the way to New Zealand for more data.” It goes without saying that both are right and the ITC Conference nourishes both of these views.
On July 30—the day after the whole-day ITC Council meeting—I attended the short course delivered by Alina von Davier (an introduction to AI-based automated item generation, in which we learned how engineering principles from the industry could be applied to the design of an item factory with agentic workflows) and Duanli Yan (an introduction to automated scoring of responses to open-ended questions). What an intensive workshop! Alina and Duanli presented a Springer book on the use of AI in education that both of them edited, and to which my colleagues Laura Casanellas Luri and Manuel Souto Pico, Dr. Alina Karakanta and I contributed a chapter.
The next day I didn’t have to figure out what session I’d attend at 8:00 in the morning: I began the day with my first presentation, Feasibility and validation study of a hybrid translation setting for PISA instruments. I was delighted to see that so many conference attendees are early birds and shared our work on evaluating the practical implications of replacing one of two human translations with a supervised automated translation, and then having (human) experts reconciling those two translations, extracting the best elements from both. We had log data from a scall-scale preliminary study showing the extent to which the reconcilers relied on each of the two translations and when.
In the same session, I was granted a glimpse into the psychometric evidence from the Duolingo English language test, learned about a model to monitor student well-being in Russian universities, and saw the results of a test of AI’s ability to answer university-level biology questions. That gives a general idea of the level and the diversity of presentations. In this session, data dominated.
Kadriye Ercikan’s presidential address gave a sober overview of the opportunities and risks of integrating AI into measurement and called for a paradigm shift. If the efficacy and the validity of AI applications cannot be critically measured, the claims about their benefits will remain unsubstantiated, and that is not acceptable for assessment practitioners. The second keynote, by Xiaoming Xi, was a perfect segue: she argued that tests should include the measurement of higher-order skills such as leveraging AI to solve complex problems. This presentation was an eye-opener for me and for the colleagues sitting close to me: the arguments were laid out in a clear and concise way, with thought-provoking diagrams and eloquence. Less data, more insights.
After lunch, we (several ITC Council members) met the Graduate students and asked them for suggestions about actions we could take (and they could take) to make the ITC more attractive to graduate students.
My second presentation was scheduled in the afternoon: much credit goes to Pavlos Stampoulidis, a psychometrician and product manager at Epignosis, who couldn’t join me in Auckland but did the lion’s share of the work behind the experiment on which I reported: we organised a blind comparison by linguists and by SMEs of a professional human translation versus LLM-assisted translation and cultural adaptation of IPIP personality assessment items from English into Greek and Indonesian. The title of the presentation was Can Machines Understand Culture? Evaluating LLM-Assisted Personality Assessment Translation Beyond WEIRD Contexts. WEIRD as in Western, Educated, Industrialised, Rich and Democratic. This was very well attended and received. And there was a wealth of data, of course.
Data was not at the heart of the discussions at the welcome reception, of course. Māori culture was, and the connection of the Māori people to the land. After the conference, I went to Waitingi to learn about Te Tiriti o Waitangi, the Treaty signed in 1840 which is a textbook example of the importance of translation: the English version uses “sovereignty”, while the Māori version uses “governance”. This led to very different interpretations. It’s a pity that cApStAn Linguistic Quality Control didn’t exist yet.
AI was a hot topic, too, of course. In a somewhat nuanced way: the ITC is not a place for a debate between illiterate AI enthusiasts and luddites, but between people who have used a variety of models in a variety of settings, have done a lot of settings and have recognised and reported flaws as well as successes. I believe “caveat” was one of the most widely used terms.
July 2nd, first thing in the morning: a symposium dubbed Emerging Challenges in the Use of Artificial Intelligence in Educational Testing & Certification, with former ITC President Steve Sireci as the discussant and as panellists the current ITC President Kadriye Ercikan-Alper, Guillermo Solano-Flores, Professor at Stanford, Isabelle Gonthier, Chief Assessment Officer at PSI & ETS, Sergio Araneda, Senior Researcher at Caveon who convened the panel, and yours truly to question whether the promises of AI will convert into less inequality and to challenge the dominance of the WEIRD world view as inherent bias in too many foundation models. Less data, more ethical concerns, more philosophical considerations.
The third keynote, by Professor Elizabeth McKinley, did posit that assessments are largely cultural products of a society and touched on the necessary and controversial topic of decolonisation applied to assessments.
From the fourth keynote, by Professor Hua-hua Chang, I took away the metaphor of a floodlight and a projector to explain that, in computer adaptive testing, it is preferable to have items with low discrimination at the beginning of the test (a floodlight with a broad light beam) and items with high discrimination at later stages of the test (a projector with a sharp, narrow beam to focus on specific capabilities).
More sessions, each one intriguing, difficult choices had to be made as the scientific programme was dense and elaborate. cApStAn sponsored one of the coffee breaks on July 2nd and a second one on July 3rd, but between those two days there was an unforgettable gala dinner at the Lula Inn on the water’s edge, at the Auckland’s Princes Wharf. Fresh seafood to share, pies, mutton, everything was fresh, the vibe was upbeat, and we had ample time to catch up with our peers and old friends, and also to make new acquaintances in the assessment ecosystem. In some ways, a gala dinner is an agreeable of a low-stakes socialising test. All the people I talked to passed with flying colours.
There was a fill house again at the first symposium of the next day: AI and Test Validity: New Frontiers in Test Development and Interpretation, with Javier Suárez-Álvarez from UMass Amherst as chair and Alina von Davier as discussant. Here, the data prevailed again, and the quality of each presentation was remarkable.
So, more data or more responses to ethical concerns? I believe that in the end a consensus was reached to regard both as indispensable and that both were equally addressed at this edition of the ITC. A big thank you to Professor Gavin Brown from the University of Auckland and the entire organising team at UoA for a flawless organisation in an extraordinary venue.