Empowering UX researchers with a plug-and-play platform to design, set up, and test human-agent interaction
The design and development of a modular AI-testing platform from scratch, leveraging an AI-powered design stack.
The challenge
With AI agents becoming deeply integrated into daily life, UX researchers and designers face a growing need to understand how people interact with and place trust in these systems. Creating experiments or running A/B tests across varying AI agents, however, often demands technical expertise or developer assistance.
To address these technical barriers, we are developing ASTRA: a platform that empowers researchers to seamlessly design, configure, and evaluate human-agent interactions.
- My role
- Lead product designer
- Partners
- imec (owner)
- Status
- Ongoing
The process
01 — Ideation
From prompt to concept
To bring this concept of a human-agent interaction testing platform to life, Claude Code and Google Stitch were leveraged to co-create the platform’s initial interface design.
To ensure high-quality outputs from the LLMs, the prompts strictly adhered to the RTCCF framework (Role, Task, Context, Constraints, and Format). This AI-driven ideation phase allowed us to rapidly map out and visualise our very first conceptual model.
02 — Design & development
Leveraging AI tools to rapidly prototype the initial proof of concept
Once the concept and style were set, an AI-powered design stack was leveraged to turn the static designs into a functional proof of concept in just a few weeks.
The AI-powered stack
03 — The proof of concept
A fully functional version ready for testing
The proof of concept allows researchers to set up human-agent experiments in just 4 simple steps.
The researcher sets up the agent
By simply toggling settings on and off, the researcher can select the agent’s AI model, set up its capabilities and tools (e.g., live news access), define its visual interface (chat, voice, or avatar), and adjust behavioural settings like tone of voice.
The researcher chooses what to measure
Built-in sensors within ASTRA track participant interactions with the agent, offering options ranging from basic questionnaires and text sentiment analysis to advanced tools like facial emotion recognition and eye tracking.
The participant interacts with the agent
On their screen, the participant interacts naturally by typing, speaking, or engaging face-to-face with the avatar, while the platform records real-time behavioural, emotional, and interaction metrics.
The researcher monitors the interaction live
Researchers can follow interactions in real time from their own device via a live dashboard, which features a synchronised timeline of messages alongside sensor data like facial emotions and eye tracking. Sessions can also be easily exported as a clean CSV for further analysis.
Agent Arena mode
For researchers conducting A/B testing between two distinct agent configurations, the Agent Arena mode enables side-by-side comparison, allowing participants to compare setups directly and express their preference.
04 — User testing
Conducting real-world trials with VRT
To evaluate initial platform usability, we partnered with public broadcaster VRT and VUB (Vrije Universiteit Brussel) to compare their custom agent, designed to answer news questions using the VRT archives, against general-purpose AI models. We configured an A/B test within the Agent Arena and conducted an initial trial with 6 participants.
This initial trial successfully demonstrated ASTRA’s real-world capabilities, generating strong enthusiasm among participants and researchers alike. The collected feedback was structured in an affinity diagram to directly guide and inspire the next iteration of the platform.
05 — Impact