JSALT 2026 research teams explore multimodal AI, conversational systems, and trustworthy health coaching

After completing an intensive two-week summer school, participants in the 12th Frederick Jelinek Memorial Summer Workshop on Speech and Language Technology (JSALT 2026) have transitioned into six weeks of collaborative research aimed at advancing the future of speech, language, and…

After completing an intensive two-week summer school, participants in the 12th Frederick Jelinek Memorial Summer Workshop on Speech and Language Technology (JSALT 2026) have transitioned into six weeks of collaborative research aimed at advancing the future of speech, language, and artificial intelligence.

Hosted by the Center for Language and Speech Processing (CLSP), the eight-week program brings together more than 50 researchers and students from institutions around the world to tackle ambitious challenges in speech and language technology. This year’s three interdisciplinary research teams are advancing several projects, including:

multimodal AI, which integrates many types of data, such as text, images and audio, at once,
spoken conversational systems, such as text-to-speech and chatbots,
and uncertainty-aware health coaching, which uses machine learning tools that acknowledge when the model doesn’t have enough information to provide guidance.
Omnimodal Encoders: When AI Goes Beyond Text

One of the projects, Omnimodal Encoders, is exploring how large language models can more effectively process information beyond text, including images, speech, video, and environmental sounds.

The researchers are developing an evaluation framework that can identify which encoder models are most likely to perform well in multimodal AI systems, which combine text, images, speech, and other types of data, without requiring the costly process of training a full model from scratch.

According to team lead David Harwath, one of the group’s biggest challenges has been selecting a diverse yet manageable set of benchmarks from the rapidly expanding number of available datasets and evaluation methods. The team also worked with JSALT organizers and sponsors to secure the computational resources needed to support its research goals.

By the end of the workshop, the researchers hope to demonstrate the effectiveness of their evaluation framework and develop new strategies for building improved multimodal encoders.

For student participant Sandra Arcos-Holzinger, a PhD candidate at the University of Melbourne and a visiting graduate scholar at Johns Hopkins University, the workshop’s collaborative environment has been one of the highlights of the experience. “Participating in JSALT is a great opportunity to collaborate with peers and experts from various institutions,” she said. “Everyone brings so much knowledge across domains that I’m excited to see what we uncover to better understand robustness in multimodal and omnimodal encoders.”

Tackling the Uncertainty in Health Coaching

Another team, Uncertainty-Aware Health Coaching for Sustainable Habit Building with Human Oversight, is focused on creating AI systems that can support healthy behavior change while remaining transparent about uncertainty and maintaining meaningful human involvement.

The team has established the project’s core components, including a conversational agent, which replicates natural dialogue, an uncertainty module that helps AI systems avoid speculations when the system lacks information, a simulated user agent that replicates how humans interact with AI, and an evaluation framework designed to assess both general system performance and health-coaching-specific behaviors.

 

 

Researchers say their greatest challenge has been integrating these interconnected research directions into a single, cohesive system. They addressed this by organizing the work around clearly defined components and interfaces while maintaining close collaboration across the team.

By the end of the workshop, the researchers aim to deliver a working end-to-end prototype for uncertainty-aware health coaching, along with methods for evaluating safety, coaching quality, and uncertainty handling in AI-assisted interventions.

For Angus Qi Chwen Ong, a PhD research student working on health-related large language models, the workshop has demonstrated the value of interdisciplinary collaboration. “One of the most rewarding aspects of the JSALT workshop has been collaborating with researchers from both academia and industry, bringing together deep domain and technical expertise,” he says. “The team’s highly dynamic structure has created a collaborative environment where diverse perspectives are genuinely valued. Working in this fast-paced, open setting has challenged me to think beyond my own discipline and approach problems from multiple angles.”

Refining Chatbots

A third project, Simulating Full-Duplex Conversations for Evaluating AI Systems, is developing new ways to evaluate spoken conversational AI by creating a realistic user simulator capable of interacting naturally with AI systems.

As conversational AI becomes increasingly sophisticated, users expect to interrupt responses, provide backchannel cues, express emotion, and engage in the fluid exchanges that characterize human conversation. Existing evaluation methods, however, often rely on simplified simulations that cannot fully capture these natural interactions.

 

The team aims to build a simulator capable of generating realistic spoken conversations while enabling systematic, repeatable evaluation of conversational AI systems. The project focuses on developing user simulation, collaborative human-AI data annotation, and evaluation methods that measure conversational quality, competence, and paralinguistic behaviors such as emotion and speaking style.