🎓 On the job market. Defending my Ph.D. thesis in August 2026 and graduating in November 2026. I’m looking for research positions in HCI and human–agent interaction, in academia and industry labs alike.
My name is Dingdong LIU. I’m currently a Ph.D. candidate in the Department of Computer Science and Engineering (CSE) at the HCI Initiative of HKUST, supervised by Prof. Xiaojuan MA and Prof. Fugee TSUNG since 2022. I also do research on humanoid robots with Prof. Bertram Emil SHI at the Center for Aging Science. Prior to that, I earned my Bachelor’s degree in Computer Science and Data Science at HKUST.
My work asks how conversational agents can acquire communicative competence — the ability to hold a conversation the way a skilled human interlocutor would. My thesis develops this at two levels:
Coordination — reading conversational dynamics: when to speak, how to signal and yield the floor, and how to handle interruptions. I study this in embodied spoken dialogue systems and humanoid robots.
Content — deciding what the conversation should be about. I approach this two ways: empirically, by learning from and co-designing with domain experts, and by inference, using human modeling and graph algorithms to derive the next move.
Together these ask what an agent must be able to do to genuinely support people — in clinical interviews, cognitive screening, and everyday decision making.
Research interests: Human–Agent Interaction · Conversational Agents · Human–Robot Interaction · Turn-Taking & Dialogue Coordination · Health & Clinical HCI · Large Language Models
CoNarrate was conditionally accepted to UIST 2026 as first author — on modeling information prematureness and scaffolding patient narratives in asynchronous telehealth.
May 04, 2026
Received the HKUST RedBird Academic Excellence Award for Continuing PhD Students (2025-26).
Visiting the Pervasive HCI Group at Tsinghua University, hosted by Prof. Yuntao Wang, to work on clinician-patient-agent collaboration in post-dentist care.
An independent project to improve telehealth communication, analyzing 5,000+ encounters and resulting in a taxonomy of information-delivery failures. Developed a conversational agent that significantly accelerated communication and improved information accuracy and relevance in patient narratives.
Under Review
Dissecting Interaction Gulfs: Towards Automating Neurocognitive Disorder Screening Dialog Tasks
Dingdong Liu, Lueyang Zhang, Junze Li, Hei Ching Iu, Bolin Zhao, Ka Ho Wong, Helen Meng, Xiaojuan Ma, and Jiaxiong Hu
Deploys a conversational agent combining coordination-level and content-level communicative abilities in a neurocognitive disorder screening scenario, where older adults must make sense of an unfamiliar and complex system. Using the gulfs of execution and evaluation to analyze interactions with 38 older adults, we identify the interaction gulfs that can undermine test validity and derive design considerations for automating screening dialog tasks.
Hospital admission interviews are critical for patient care but strain nurses’ capacity due to time constraints and staffing shortages. While LLM-powered conversational agents (CAs) offer automation potential, their rigid sequencing and lack of humanized communication skills risk misunderstandings and incomplete data capture. Through participatory design with clinicians and volunteers, we identified essential communication strategies and developed a novel CA that implements these strategies through: (1) dynamic topic management using graph-based conversation flows, and (2) context-aware scaffolding with few-shot prompt tuning. Technical evaluation on an admission interview dataset showed our system achieving performance comparable to or surpassing human-written ground truth, while outperforming prompt-engineered baselines. A between-subject study (N=44) demonstrated significantly improved user experience and data collection accuracy compared to existing solutions. We contribute a framework for humanizing medical CAs by translating clinician expertise into algorithmic strategies, alongside empirical insights for balancing efficiency and empathy in healthcare interactions, and considerations for generalizability.
@inproceedings{liuScaffoldedTurnsAndLogicalConversationsDesigning2025,author={Liu, Dingdong and Zhang, Yujing and Zhao, Bolin and Ma, Shuai and Shi, Chuhan and Ma, Xiaojuan},title={Scaffolded Turns and Logical Conversations: Designing Humanized LLM-Powered Conversational Agents for Hospital Admission Interviews},year={2025},isbn={9798400713941},publisher={Association for Computing Machinery},address={New York, NY, USA},url={https://doi.org/10.1145/3706598.3714196},doi={10.1145/3706598.3714196},booktitle={Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems},articleno={643},numpages={23},keywords={Conversational Agents, Clinical Communication, Hospital Admission Interview, Healthcare Automation, Scaffolded Dialogue, Participatory Design, Large Language Models},series={CHI '25}}
The ability to coordinate turn taking during spoken dialogue is crucial for an embodied spoken dialogue system (SDS), e.g., in a humanoid robot. The SDS needs to model transitions in the conversational floor, which describes each party’s stance (either speaking or listening). Further, the SDS needs to signal its perception of the floor to the human, so that they can coordinate floor transitions and resolve conflicts. Conventional SDS employ standalone modules to control floor transitions but do not produce timely and appropriate responses. Recent end-to-end audio LLMs generate responses quickly, but do not coordinate floor transitions as accurately. In this work, we propose an SDS architecture that dynamically adjusts its prompts to an end-to-end audio LLM based upon its perception of the conversational floor state. The LLM output determines not only the audio output, but also the perceived floor state. This enables the system to signal its stance to the human, both when listening and when speaking. We conducted an experiment where a humanoid robot administered a semi-structured interview with human subjects. Results show that, compared with baseline systems using static prompts, dynamic prompting enables the LLM to model floor transitions more accurately, to generate more appropriate signalling, and to interrupt less, leading to smoother turn-taking in dialogue.
@inproceedings{shenDynamicPromptingImproves2025,author={Shen, Yifan and Liu, Dingdong and Mo, Xiaoyu and Tsung, Fugee and Ma, Xiaojuan and Shi, Bertram E.},booktitle={2025 34th IEEE International Conference on Robot and Human Interactive Communication (RO-MAN)},title={Dynamic Prompting Improves Turn-taking in Embodied Spoken Dialogue Systems},year={2025},pages={692-699},keywords={Navigation;Robot kinematics;Humanoid robots;Systems architecture;Oral communication;Manuals;History;Interviews;Floors;Signal resolution},doi={10.1109/RO-MAN63969.2025.11217726}}
🏆 A Humanoid Robot Dialogue System Architecture Targeting Patient Interview Tasks
Yifan Shen, Dingdong Liu, Yejin Bang, Ho Shu Chan, Rita Frieske, Hoo Choun Chung, Jay Nieles, Tianjia Zhang, Kien T. Pham, Wai Yi Rosita Cheng, Yini Fang, Qifeng Chen, Pascale Fung, Xiaojuan Ma, and Bertram E. Shi
In 2024 33rd IEEE International Conference on Robot and Human Interactive Communication (ROMAN). Co-first author, equal contribution , 2024
UBTECH Best Industry Application Award (Bronze), IEEE RO-MAN 2024.
Humanoid robots are promising approach to automating patient interviews routinely conducted by medical staff. Their human-like appearance enables them to use the full gamut of verbal and behavioral cues that are critical to a successful interview. On the other hand, anthropomorphism can induce expectations of human-level performance by the robot. Not meeting such expectations degrades the quality of interaction. Specifically, humans expect rich real-time interactions during speech exchange, such as backchanneling and barge-ins. The nature of the patient interview task differs from most other scenarios where task oriented dialogue systems have been used, as there is increased potential of engagement breakdown during interaction. We describe a dialogue system architecture that improves the performance of humanoid robots on the patient interview task. Our architecture adds a nested inner real-time control loop to improve the timeliness of the robot’s responses based on the notion of "stance", an elaboration of the concept of a "turn", common in most existing dialogue systems. It also expands the dialogue state to monitor not only task progress, but also human engagement. Experiments using a humanoid robot running our proposed architecture reveal improved performance on interview tasks in terms of the perceived timeliness of responses and users’ impressions of the system.
@inproceedings{shenHumanoidRobotDialogueSystemArchitecture2024,award_photo={roman2024-award-photo.jpg},author={Shen, Yifan and Liu, Dingdong and Bang, Yejin and Chan, Ho Shu and Frieske, Rita and Chung, Hoo Choun and Nieles, Jay and Zhang, Tianjia and Pham, Kien T. and Cheng, Wai Yi Rosita and Fang, Yini and Chen, Qifeng and Fung, Pascale and Ma, Xiaojuan and Shi, Bertram E.},booktitle={2024 33rd IEEE International Conference on Robot and Human Interactive Communication (ROMAN)},title={A Humanoid Robot Dialogue System Architecture Targeting Patient Interview Tasks},year={2024},pages={1394-1401},keywords={Target tracking;Navigation;Robot kinematics;Humanoid robots;Systems architecture;Real-time systems;History;Interviews;Standards;Monitoring},doi={10.1109/RO-MAN60168.2024.10731285}}
Cognitive screening in hospitalized older patients is a critical, yet time-consuming process. While conversational agents present a promising solution to aid clinicians, current models fall short in their ability to scaffold questions to accommodate patients with potential cognitive decline effectively. To bridge this gap, we conducted a study with 13 clinicians to identify effective scaffolding strategies empirically. Our findings revealed six key strategies that clinicians use to scaffold the Abbreviated Mental Test (AMT) in practice, together with the underlying rationale and potential challenges. We discuss the implications of these findings for the design of conversational agents to assist in cognitive screening and propose design considerations for future research.
@inproceedings{liuExploringScaffoldingTechniques2024,author={Liu, Dingdong and Gao, Sensen and Chen, Zixin and Shen, Yifan and Shi, Chuhan and Shi, Bertram E. and Ma, Xiaojuan},title={Exploring Scaffolding Techniques for Agent-Administered Brief Cognitive Screening in Hospital Settings},year={2024},isbn={9798400706325},publisher={Association for Computing Machinery},address={New York, NY, USA},url={https://doi.org/10.1145/3656156.3663697},doi={10.1145/3656156.3663697},booktitle={Companion Publication of the 2024 ACM Designing Interactive Systems Conference},pages={185–189},numpages={5},keywords={cognitive screening, conversational agent, scaffolding},location={IT University of Copenhagen, Denmark},series={DIS '24 Companion}}