Publications研究成果

RESEARCH OUTPUT · 40+ PUBLICATIONS研究成果 · 40 余篇论文

Jiangtao Gong has published 40+ papers in leading HCI, AI, and robotics venues, including CHI, UIST, CSCW, NeurIPS, AAAI, ICRA, IROS, WWW, ICCV, and ICLR. See the complete record on Google Scholar.

龚江涛以第一作者或通讯作者身份,在 CHI、UIST、CSCW、NeurIPS、AAAI、ICRA、IROS、WWW、ICCV、ICLR 等人机交互、人工智能与机器人领域重要会议和期刊发表论文 40 余篇。完整列表请见 Google Scholar

Intelligent System Design智能系统设计54
Interaction Design交互设计18

An LLM-based simulation framework for embodied conversational agents in psychological counseling Permalink

Published in Proceedings of the AAAI Conference on Artificial Intelligence 40 (35), 29758 ..., 2026

Authors: L Wu, Y Tang, Q Pan, X Zhan, Y Han, L Xiao, T Wang, C Zhong, Jiangtao Gong

Abstract: Due to privacy concerns, open dialogue datasets for mental health are primarily generated through human or AI synthesis methods. However, the inherent implicit nature of psychological processes, particularly those of clients, poses challenges to the authenticity and diversity of synthetic data. In this paper, we propose ECAs (short for Embodied Conversational Agents), a framework for embodied agent simulation based on Large Language Models (LLMs) that incorporates multiple psychological theoretical principles. Using simulation, we expand real counseling case data into a nuanced embodied cognitive memory space and generate dialogue data based on high-frequency counseling questions. We validated our framework using the D4 dataset. First, we created a public ECAs dataset through batch simulations based on D4. Licensed counselors evaluated our method, demonstrating that it significantly outperforms baselines in simulation authenticity and necessity. Additionally, two LLM-based automated evaluation methods were employed to confirm the higher quality of the generated dialogues compared to the baselines.

Recommended citation: L Wu, Y Tang, Q Pan, X Zhan, Y Han, L Xiao, T Wang, C Zhong, J Gong (2026). An LLM-based simulation framework for embodied conversational agents in psychological counseling. Proceedings of the AAAI Conference on Artificial Intelligence 40 (35), 29758 ....

Mentigo: An intelligent agent for mentoring students in the creative problem solving process Permalink

Published in Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems ..., 2025

Authors: S Zha, Y Liu, C Zheng, J Xu, F Yu, Jiangtao Gong, Y Xu

Abstract: Creative Problem-Solving (CPS) promotes creative and critical thinking while enhancing real-world problem-solving skills, making it essential for middle school education.However, providing personalized mentorship in CPS projects at scale is challenging due to resource constraints and diverse student needs.To address this, we developed Mentigo, an AI-driven mentor agent designed to guide middle school students through the CPS process.Using a dataset of real classroom interactions, we encoded CPS task stages, adaptive guidance strategies, and personalized feedback mechanisms to inform Mentigo's dynamic mentoring framework powered by large language models (LLMs).A comparative experiment with 12 students and evaluations from five expert educators demonstrated improved student engagement, creativity, and task performance.Our findings highlight design implications for using LLM-based AI mentors to enhance CPS learning in educational environments.

Recommended citation: S Zha, Y Liu, C Zheng, J Xu, F Yu, J Gong, Y Xu (2025). Mentigo: An intelligent agent for mentoring students in the creative problem solving process. Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems ....

Designing LLM-simulated Immersive Spaces to Enhance Autistic Children’s Social Affordances Understanding in Traffic Settings Permalink

Published in Proceedings of the 30th International Conference on Intelligent User ..., 2025

Authors: Y Cao, Y He, Y Chen, M Chen, S You, Y Qiu, M Liu, C Luo, C Zheng, …, Jiangtao Gong

Abstract: One of the key challenges faced by autistic children is understanding social affordances in complex environments, which further impacts their ability to respond appropriately to social signals. In traffic scenarios, this impairment can even lead to safety concerns. In this paper, we introduce an LLM-simulated immersive projection environment designed to improve this ability in autistic children while ensuring their safety. We first propose 17 design considerations across four major categories, derived from a comprehensive review of previous research. Next, we developed a system called AIroad, which leverages LLMs to simulate drivers with varying social intents, expressed through explicit multimodal social signals. AIroad helps autistic children bridge the gap in recognizing the intentions behind behaviors and learning appropriate responses through various stimuli. A user study involving 14 participants demonstrated that this technology effectively engages autistic children and leads to significant improvements in their comprehension of social affordances in traffic scenarios. Additionally, parents reported high perceived usability of the system. These findings highlight the potential of combining LLM technology with immersive environments for the functional rehabilitation of autistic children in the future.

Recommended citation: Y Cao, Y He, Y Chen, M Chen, S You, Y Qiu, M Liu, C Luo, C Zheng,... (2025). Designing LLM-simulated Immersive Spaces to Enhance Autistic Children's Social Affordances Understanding in Traffic Settings. Proceedings of the 30th International Conference on Intelligent User ....

Designing child-centric AI learning environments: Insights from an LLM-powered creative project-based learning study Permalink

Published in International Journal of Human-Computer Studies, 103602, 2025

Authors: S Zha, Y Qiao, Q Hu, Z Li, Jiangtao Gong, Y Xu

Abstract: Project-based learning (PBL) is an instructional method that is very helpful in nurturing students' creativity, but it requires significant time and energy from both students and teachers. Large language models (LLMs) have been proven to assist in creative tasks, yet much controversy exists regarding their role in fostering creativity. This paper explores the potential of LLMs in PBL settings, with a special focus on fostering creativity. We began with an exploratory study involving 12 middle school students and identified five design considerations for LLM applications in PBL. Building on this, we developed an LLM-empowered, 48-hour PBL program and conducted an instructional experiment with 31 middle school students. Our results indicated that LLMs can enhance every stage of PBL. Additionally, we also discovered ambivalent perspectives among students and mentors toward LLM usage. Furthermore, we explored the challenge and design implications of integrating LLMs into PBL and reflected on the program. By bridging AI advancements into educational practice, our work aims to inspire further discourse and investigation into harnessing AI's potential in child-centric educational settings.

Recommended citation: S Zha, Y Qiao, Q Hu, Z Li, J Gong, Y Xu (2025). Designing child-centric AI learning environments: Insights from an LLM-powered creative project-based learning study. International Journal of Human-Computer Studies, 103602.

COLP: scaffolding children’s online long-term collaborative learning Permalink

Published in International Journal of Human-Computer Interaction 41 (24), 15453-15475, 2025

Authors: S Zha, Y Tang, Jiangtao Gong, Y Xu

Abstract: Online collaborative learning and working are important for everyone including children. However, children still face a lot of difficulties communicating and working together while online, which keeps them from engaging in long-term project-based teamwork. We aim to investigate online long-term collaborative learning opportunities to address this gap. We design COLP, an online, 16-week, project-based learning program, as an educational intervention based on multiple learning theories for primary school students. We conducted this program with 67 primary school students ages 8-13, across more than five provinces of China. We found that this program could engage more than one-third of children in teamwork after long-term study. Furthermore, we interview children and their parents to help us understand the communication channel, benefits, and challenges of this program. Interestingly, we discovered that parents play multiple roles in their children's collaborative learning, particularly modeling and guiding the children's collaborative skills. Given the lack of programs designed for children's long-term online collaboration, this study may inspire intervention design in computer-supported collaborative learning communities.

Recommended citation: S Zha, Y Tang, J Gong, Y Xu (2025). COLP: scaffolding children's online long-term collaborative learning. International Journal of Human-Computer Interaction 41 (24), 15453-15475.

AI system facilitates people with blindness and low vision in interpreting and experiencing unfamiliar environments Permalink

Published in npj Artificial Intelligence 1 (1), 7, 2025

Authors: H Lin, Jiangtao Gong, Y Wang, J Zhang, B Bai, Y Zhang, L Wang, C Wei, Y Cao, …

Abstract: Engaging with nature significantly enhances well-being, yet millions of individuals with blindness and low vision (BLV) are often excluded from these benefits due to constrained environmental perception. Here, we introduce VIPTour, an AI-driven system powered by the FocusFormer algorithm, which transforms complex scenes into structured, personalized graphs using tailored attention mechanisms and a BLV-in-the-Loop Adapter. Through intuitive, user-centered interaction, VIPTour facilitates active exploration and in-depth comprehension during dynamic sightseeing, while enabling accurate, long-lasting recollection and effective communication among BLV individuals post-journey. Extensive experiments demonstrate that VIPTour significantly enhances positive emotions and memory retention, with a 67.9% increase in positive emotional response, a 94.7% rise in arousal, a 772.73% improvement in cognitive mapping accuracy, and a 200% enhancement in long-term memory accuracy. These results underscore VIPTour’s ability to deliver an unparalleled, enjoyable, and memorable experience, promising profound benefits for the BLV community.

Recommended citation: H Lin, J Gong, Y Wang, J Zhang, B Bai, Y Zhang, L Wang, C Wei, Y Cao,... (2025). AI system facilitates people with blindness and low vision in interpreting and experiencing unfamiliar environments. npj Artificial Intelligence 1 (1), 7.

Guiding Multiple Remote Users in Physical Tasks with Language-driven Robotic Telepresence Permalink

Published in Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous ..., 2025

Authors: R Li, J Guo, X Zhang, X Zhang, Z Li, J Li, Jiangtao Gong

Abstract: Remote assistance through robotic telepresence could involve both control and memory challenges, particularly in one expert to multiple workers situation. In this work, we proposed a novelty language-driven interface to facilitate remote collaboration through telepresence robots. Through operations and maintenance expert interviews and a scenario simulation study, we identified key pain points in executing one-expert-multiple-workers remote guidance using the telepresence robot and proposed two design goals, which together consist of five sub-design goals with corresponding features. These features were integrated into a standard telepresence robot, resulting in the development of a Collaborative LLM-based Embodied Assistant Robot, named CLEAR Robot. A controlled experiment simulating a remote assembly task of one to two demonstrated that, compared to the standard telepresence robot, CLEAR Robot significantly improved efficiency, reduced cognitive load, facilitated more balanced collaboration, and improved the user experience. We also discuss the impact of language-driven implicit interactions in multi-user collaboration and provide insights for designing robot systems that support one-expert-multiple-workers remote guidance in the future.

Recommended citation: R Li, J Guo, X Zhang, X Zhang, Z Li, J Li, J Gong (2025). Guiding Multiple Remote Users in Physical Tasks with Language-driven Robotic Telepresence. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous ....

TouchEditor: interaction design and evaluation of a flexible touchpad for text editing of head-mounted displays in speech-unfriendly environments Permalink

Published in Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous ..., 2024

Authors: L Zhan, T Xiong, H Zhang, S Guo, X Chen, Jiangtao Gong, J Lin, Y Qin

Recommended citation: L Zhan, T Xiong, H Zhang, S Guo, X Chen, J Gong, J Lin, Y Qin (2024). TouchEditor: interaction design and evaluation of a flexible touchpad for text editing of head-mounted displays in speech-unfriendly environments. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous ....

Teleaware robot: Designing awareness-augmented telepresence robot for remote collaborative locomotion Permalink

Published in Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous ..., 2024

Authors: R Li, Y Zhu, M Liu, Y Zeng, S Zhuang, J Fu, Y Lu, G Zhou, C Liu, Jiangtao Gong

Abstract: Telepresence robots can be used to support users to navigate an environment remotely and share the visiting experience with their social partners. Although such systems allow users to see and hear the remote environment and communicate with their partners via live video feed, this does not provide enough awareness of the environment and their remote partner's activities. In this paper, we introduce an awareness framework for collaborative locomotion in scenarios of onsite and remote users visiting a place together. From an observational study of small groups of people visiting exhibitions, we derived four design goals for enhancing the environmental and social awareness between social partners, and developed a set of awareness-enhancing techniques to add to a standard telepresence robot - named TeleAware robot. Through a controlled experiment simulating a guided exhibition visiting task, TeleAware robot showed the ability to lower the workload, facilitate closer social proximity, and improve mutual awareness and social presence compared with the standard one. We discuss the impact of mobility and roles of local and remote users, and provide insights for the future design of awareness-enhancing telepresence robot systems that facilitate collaborative locomotion.

Recommended citation: R Li, Y Zhu, M Liu, Y Zeng, S Zhuang, J Fu, Y Lu, G Zhou, C Liu, J Gong (2024). Teleaware robot: Designing awareness-augmented telepresence robot for remote collaborative locomotion. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous ....

“I am the follower, also the boss”: Exploring Different Levels of Autonomy and Machine Forms of Guiding Robots for the Visually Impaired Permalink

Published in Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems ..., 2023

Authors: Y Zhang, Z Li, H Guo, L Wang, Q Chen, W Jiang, M Fan, G Zhou, Jiangtao Gong

Abstract: Guiding robots, in the form of canes or cars, have recently been explored to assist blind and low vision (BLV) people. Such robots can provide full or partial autonomy when guiding. However, the pros and cons of different forms and autonomy for guiding robots remain unknown. We sought to fill this gap. We designed autonomy-switchable guiding robotic cane and car. We conducted a controlled lab-study (N=12) and a field study (N=9) on BLV. Results showed that full autonomy received better walking performance and subjective ratings in the controlled study, whereas participants used more partial autonomy in the natural environment as demanding more control. Besides, the car robot has demonstrated abilities to provide a higher sense of safety and navigation efficiency compared with the cane robot. Our findings offered empirical evidence about how the BLV community perceived different machine forms and autonomy, which can inform the design of assistive robots.

Recommended citation: Y Zhang, Z Li, H Guo, L Wang, Q Chen, W Jiang, M Fan, G Zhou, J Gong (2023). " I am the follower, also the boss": Exploring Different Levels of Autonomy and Machine Forms of Guiding Robots for the Visually Impaired. Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems ....

MR. Brick: designing a remote mixed-reality educational game system for promoting children’s social & collaborative skills Permalink

Published in Proceedings of the 2023 CHI conference on human factors in computing systems ..., 2023

Authors: Y Wu, S You, Z Guo, X Li, G Zhou, Jiangtao Gong

Abstract: Children are one of the groups most influenced by COVID-19-related social distancing, and a lack of contact with peers can limit their opportunities to develop social and collaborative skills. However, remote socialization and collaboration as an alternative approach is still a great challenge for children. This paper presents MR.Brick, a Mixed Reality (MR) educational game system that helps children adapt to remote collaboration. A controlled experimental study involving 24 children aged six to ten was conducted to compare MR.Brick with the traditional video game by measuring their social and collaborative skills and analyzing their multi-modal playing behaviours. The results showed that MR.Brick was more conducive to children’s remote collaboration experience than the traditional video game. Given the lack of training systems designed for children to collaborate remotely, this study may inspire interaction design and educational research in related fields.

Recommended citation: Y Wu, S You, Z Guo, X Li, G Zhou, J Gong (2023). MR. Brick: designing a remote mixed-reality educational game system for promoting children's social & collaborative skills. Proceedings of the 2023 CHI conference on human factors in computing systems ....

Touch-and-heal: Data-driven affective computing in tactile interaction with robotic dog Permalink

Published in Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous ..., 2023

Authors: S Guo, L Zhan, Y Cao, C Zheng, G Zhou, Jiangtao Gong

Recommended citation: S Guo, L Zhan, Y Cao, C Zheng, G Zhou, J Gong (2023). Touch-and-heal: Data-driven affective computing in tactile interaction with robotic dog. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous ....

Enable natural tactile interaction for robot dog based on large-format distributed flexible pressure sensors Permalink

Published in arXiv preprint arXiv:2303.07595, 2023

Authors: L Zhan, Y Cao, Q Chen, H Guo, J Gao, Y Luo, S Guo, G Zhou, Jiangtao Gong

Abstract: Touch is an important channel for human-robot interaction, while it is challenging for robots to recognize human touch accurately and make appropriate responses. In this paper, we design and implement a set of large-format distributed flexible pressure sensors on a robot dog to enable natural human-robot tactile interaction. Through a heuristic study, we sorted out 81 tactile gestures commonly used when humans interact with real dogs and 44 dog reactions. A gesture classification algorithm based on ResNet is proposed to recognize these 81 human gestures, and the classification accuracy reaches 98.7%. In addition, an action prediction algorithm based on Transformer is proposed to predict dog actions from human gestures, reaching a 1-gram BLEU score of 0.87. Finally, we compare the tactile interaction with the voice interaction during a freedom human-robot-dog interactive playing study. The results show that tactile interaction plays a more significant role in alleviating user anxiety, stimulating user excitement and improving the acceptability of robot dogs.

Recommended citation: L Zhan, Y Cao, Q Chen, H Guo, J Gao, Y Luo, S Guo, G Zhou, J Gong (2023). Enable natural tactile interaction for robot dog based on large-format distributed flexible pressure sensors. arXiv preprint arXiv:2303.07595.

Holoboard: A large-format immersive teaching board based on pseudo holographics Permalink

Published in The 34th Annual ACM Symposium on User Interface Software and Technology, 441-456, 2021

Authors: Jiangtao Gong, T Han, S Guo, J Li, S Zha, L Zhang, F Tian, Q Wang, Y Rui

Abstract: In this paper, we present HoloBoard, an interactive large-format pseduo-holographic display system for lecture based classes. With its unique properties of immersive visual display and transparent screen, we designed and implemented a rich set of novel interaction techniques like immersive presentation, role-play, and lecturing behind the scene that are potentially valuable for lecturing in class. We conducted a controlled experimental study to compare a HoloBoard class with a normal class through measuring students’ learning outcomes and three dimensions of engagement (i.e., behavioral, emotional, and cognitive engagement). We used pre-/post- knowledge tests and multimodal learning analytics to measure students’ learning outcomes and learning experiences. Results indicated that the lecture-based class utilizing HoloBoard lead to slightly better learning outcomes and a significantly higher level of student engagement. Given the results, we discussed the impact of HoloBoard as an immersive media in the classroom setting and suggest several design implications for deploying HoloBoard in immersive teaching practices.

Recommended citation: J Gong, T Han, S Guo, J Li, S Zha, L Zhang, F Tian, Q Wang, Y Rui (2021). Holoboard: A large-format immersive teaching board based on pseudo holographics. The 34th Annual ACM Symposium on User Interface Software and Technology, 441-456.

HeliCoach: 基于无人机模拟 3D 音频空间的多通道自适应定向行走训练系统的设计研究 Permalink

Published in 计算机辅助设计与图形学学报 32 (7), 1129-1136, 2020

Authors: 龚江涛,丁琦城,徐鹏辉,张宇,张柳新,王茜莺

Recommended citation: 龚江涛,丁琦城,徐鹏辉,张宇,张柳新,王茜莺 (2020). HeliCoach: 基于无人机模拟 3D 音频空间的多通道自适应定向行走训练系统的设计研究. 计算机辅助设计与图形学学报 32 (7), 1129-1136.

Algorithm Optimization算法优化22

FreeAskWorld: An Interactive and Closed-Loop Simulator for Human-Centric Embodied AI Permalink

Published in Proceedings of the AAAI Conference on Artificial Intelligence 40 (2), 926-934, 2026

Authors: Y Peng, Y Pan, X He, J Yang, X Yin, H Wang, X Zheng, C Gao, Jiangtao Gong

Abstract: As embodied intelligence emerges as a core frontier in artificial intelligence research, simulation platforms must evolve beyond low-level physical interactions to capture complex, human-centered social behaviors. We introduce FreeAskWorld, an interactive simulation framework that integrates large language models (LLMs) for high-level behavior planning and semantically grounded interaction, informed by theories of intention and social cognition. Our framework supports scalable, realistic human-agent simulations and includes a modular data generation pipeline tailored for diverse embodied tasks.To validate the framework, we extend the classic Vision-and-Language Navigation (VLN) task into a semantically enriched Direction Inquiry setting, wherein agents can actively seek and interpret navigational guidance. We present and publicly release FreeAskWorld, a large-scale benchmark dataset comprising reconstructed environments, six diverse task types, 16 core object categories, 63,429 annotated sample frames, and more than 17 hours of interaction data to support training and evaluation of embodied AI systems. We benchmark VLN models, and human participants under both open-loop and closed-loop settings. Experimental results demonstrate that models fine-tuned on FreeAskWorld outperform their original counterparts, achieving enhanced semantic understanding and interaction competency. These findings underscore the efficacy of socially grounded simulation frameworks in advancing embodied AI systems toward sophisticated high-level planning and more naturalistic human-agent interaction.

Recommended citation: Y Peng, Y Pan, X He, J Yang, X Yin, H Wang, X Zheng, C Gao, J Gong (2026). FreeAskWorld: An Interactive and Closed-Loop Simulator for Human-Centric Embodied AI. Proceedings of the AAAI Conference on Artificial Intelligence 40 (2), 926-934.

An LLM-based simulation framework for embodied conversational agents in psychological counseling Permalink

Published in Proceedings of the AAAI Conference on Artificial Intelligence 40 (35), 29758 ..., 2026

Authors: L Wu, Y Tang, Q Pan, X Zhan, Y Han, L Xiao, T Wang, C Zhong, Jiangtao Gong

Abstract: Due to privacy concerns, open dialogue datasets for mental health are primarily generated through human or AI synthesis methods. However, the inherent implicit nature of psychological processes, particularly those of clients, poses challenges to the authenticity and diversity of synthetic data. In this paper, we propose ECAs (short for Embodied Conversational Agents), a framework for embodied agent simulation based on Large Language Models (LLMs) that incorporates multiple psychological theoretical principles. Using simulation, we expand real counseling case data into a nuanced embodied cognitive memory space and generate dialogue data based on high-frequency counseling questions. We validated our framework using the D4 dataset. First, we created a public ECAs dataset through batch simulations based on D4. Licensed counselors evaluated our method, demonstrating that it significantly outperforms baselines in simulation authenticity and necessity. Additionally, two LLM-based automated evaluation methods were employed to confirm the higher quality of the generated dialogues compared to the baselines.

Recommended citation: L Wu, Y Tang, Q Pan, X Zhan, Y Han, L Xiao, T Wang, C Zhong, J Gong (2026). An LLM-based simulation framework for embodied conversational agents in psychological counseling. Proceedings of the AAAI Conference on Artificial Intelligence 40 (35), 29758 ....

Embodied Cognition Augmented End2End Autonomous Driving Permalink

Published in Advances in Neural Information Processing Systems 38, 44665-44688, 2026

Authors: L Niu, X Zheng, Z Yang, C Zheng, B Chen, Jiangtao Gong

Abstract: In recent years, vision-based end-to-end autonomous driving has emerged as a new paradigm. However, popular end-to-end approaches typically rely on visual feature extraction networks trained under label supervision. This limited supervision framework restricts the generality and applicability of driving models. In this paper, we propose a novel paradigm termed $E^{3}AD$, which advocates for comparative learning between visual feature extraction networks and the general EEG large model, in order to learn latent human driving cognition for enhancing end-to-end planning. In this work, we collected a cognitive dataset for the mentioned contrastive learning process. Subsequently, we investigated the methods and potential mechanisms for enhancing end-to-end planning with human driving cognition, using popular driving models as baselines on publicly available autonomous driving datasets. Both open-loop and closed-loop tests are conducted for a comprehensive evaluation of planning performance. Experimental results demonstrate that the $E^{3}AD$ paradigm significantly enhances the end-to-end planning performance of baseline models. Ablation studies further validate the contribution of driving cognition and the effectiveness of comparative learning process. To the best of our knowledge, this is the first work to integrate human driving cognition for improving end-to-end autonomous driving planning. It represents an initial attempt to incorporate embodied cognitive data into end-to-end autonomous driving, providing valuable insights for future brain-inspired autonomous driving systems. Our code will be made available at Github

Recommended citation: L Niu, X Zheng, Z Yang, C Zheng, B Chen, J Gong (2026). Embodied Cognition Augmented End2End Autonomous Driving. Advances in Neural Information Processing Systems 38, 44665-44688.

GroupRAG: Cognitively Inspired Group-Aware Retrieval and Reasoning via Knowledge-Driven Problem Structuring Permalink

Published in arXiv preprint arXiv:2603.26807, 2026

Authors: X Duan, Y Tang, Jiangtao Gong

Abstract: The performance of language models is commonly limited by insufficient knowledge and constrained reasoning. Prior approaches such as Retrieval-Augmented Generation (RAG) and Chain-of-Thought (CoT) address these issues by incorporating external knowledge or enforcing linear reasoning chains, but often degrade in real-world settings. Inspired by cognitive science, which characterizes human problem solving as search over structured problem spaces rather than single inference chains, we argue that inadequate awareness of problem structure is a key overlooked limitation. We propose GroupRAG, a cognitively inspired, group-aware retrieval and reasoning framework based on knowledge-driven keypoint grouping. GroupRAG identifies latent structural groups within a problem and performs retrieval and reasoning from multiple conceptual starting points, enabling fine-grained interaction between the two processes. Experiments on MedQA show that GroupRAG outperforms representative RAG- and CoT-based baselines. These results suggest that explicitly modeling problem structure, as inspired by human cognition, is a promising direction for robust retrieval-augmented reasoning.

Recommended citation: X Duan, Y Tang, J Gong (2026). GroupRAG: Cognitively Inspired Group-Aware Retrieval and Reasoning via Knowledge-Driven Problem Structuring. arXiv preprint arXiv:2603.26807.

A comprehensive llm-powered framework for driving intelligence evaluation Permalink

Published in 2025 IEEE International Conference on Robotics and Automation (ICRA), 7028-7035, 2025

Authors: S You, X Luo, X Liang, J Yu, C Zheng, Jiangtao Gong

Abstract: Evaluation methods for autonomous driving are crucial for algorithm optimization. However, due to the complexity of driving intelligence, there is currently no comprehensive evaluation method for the level of autonomous driving intelligence. In this paper, we propose an evaluation framework for driving behavior intelligence in complex traffic environments, aiming to fill this gap. We constructed a natural language evaluation dataset of human professional drivers and passengers through naturalistic driving experiments and post-driving behavior evaluation interviews. Based on this dataset, we developed an LLM-powered driving evaluation framework. The effectiveness of this framework was validated through simulated experiments in the CARLA urban traffic simulator and further corroborated by human assessment. Our research provides valuable insights for evaluating and designing more intelligent, human-like autonomous driving agents. The implementation details of the framework 1 1 https://github.com/AIR-DISCOVER/Driving-Intellenge-Evaluation-Framework and detailed information about the dataset 2 2 https://github.com/AIR-DISCOVER/Driving-Evaluation-Datasetcan be found at the provided links.

Recommended citation: S You, X Luo, X Liang, J Yu, C Zheng, J Gong (2025). A comprehensive llm-powered framework for driving intelligence evaluation. 2025 IEEE International Conference on Robotics and Automation (ICRA), 7028-7035.

More than routing: Joint GPS and route modeling for refine trajectory representation learning Permalink

Published in Proceedings of the ACM Web Conference 2024, 3064-3075, 2024

Authors: Z Ma, Z Tu, X Chen, Y Zhang, D Xia, G Zhou, Y Chen, Y Zheng, Jiangtao Gong

Abstract: Trajectory representation learning plays a pivotal role in supporting various downstream tasks, such as travel time estimation, trajectory classification and Top-k similar trajectory search. Traditional methods in order to filter the noise in GPS trajectories tend to focus on routing-based methods to simplify the trajectories. However, these approaches ignore the motion details contained in the GPS data, limiting the representation capability of trajectory representation learning. To fill this gap, we propose a novel representation learning framework that is Jointly G PS and Route Modeling based on self-supervised technology, namely JGRM. We consider GPS trajectory and route trajectory as the two modals of a single movement observation and fuse information through inter-modal information interaction. Specifically, we develop two encoders, each tailored to capture representations of GPS trajectories and route trajectories respectively. The representations from these two modalities are fed into a shared transformer for inter-modal information interaction. Eventually, we design three self-supervised tasks to train the model. We validate the effectiveness of the proposed method on two real-world datasets through extensive experiments. The experimental results show that JGRM significantly outperforms existing methods in both road segment representation and trajectory representation tasks. Our source code is available at Github https://github.com/mamazi0131/JGRM.

Recommended citation: Z Ma, Z Tu, X Chen, Y Zhang, D Xia, G Zhou, Y Chen, Y Zheng, J Gong (2024). More than routing: Joint GPS and route modeling for refine trajectory representation learning. Proceedings of the ACM Web Conference 2024, 3064-3075.

SurrealDriver: Designing LLM-powered generative driver agent framework based on human drivers… driving-thinking data Permalink

Published in 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems ..., 2024

Authors: Y Jin, R Yang, Z Yi, X Shen, H Peng, X Liu, J Qin, J Li, J Xie, P Gao, …, Jiangtao Gong

Abstract: Leveraging advanced reasoning capabilities and extensive world knowledge of large language models (LLMs) to construct generative agents for solving complex real-world problems is a major trend. However, LLMs inherently lack embodiment as humans, resulting in suboptimal performance in many embodied decision-making tasks. In this paper, we introduce a framework for building human-like generative driving agents using post-driving self-report driving-thinking data from human drivers as both demonstration and feedback. To capture high-quality, natural language data from drivers, we conducted urban driving experiments, recording drivers’ verbalized thoughts under various conditions to serve as chain-of-thought prompts and demonstration examples for the LLM-Agent. The framework’s effectiveness was evaluated through simulations and human assessments. Results indicate that incorporating expert demonstration data significantly reduced collision rates by 81.04% and increased human likeness by 50% compared to a baseline LLM-based agent. Our study provides insights into using natural language-based human demonstration data for embodied tasks. The driving-thinking dataset is available at https://github.com/AIR-DISCOVER/Driving-Thinking-Dataset.

Recommended citation: Y Jin, R Yang, Z Yi, X Shen, H Peng, X Liu, J Qin, J Li, J Xie, P Gao,... (2024). SurrealDriver: Designing LLM-powered generative driver agent framework based on human drivers... driving-thinking data. 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems ....

Large language models powered context-aware motion prediction in autonomous driving Permalink

Published in 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems ..., 2024

Authors: X Zheng, L Wu, Z Yan, Y Tang, H Zhao, C Zhong, B Chen, Jiangtao Gong

Abstract: Motion prediction is among the most fundamental tasks in autonomous driving. Traditional methods of motion forecasting primarily encode vector information of maps and historical trajectory data of traffic participants, lacking a comprehensive understanding of overall traffic semantics, which in turn affects the performance of prediction tasks. In this paper, we utilized Large Language Models (LLMs) to enhance the global traffic context understanding for motion prediction tasks. We first conducted systematic prompt engineering, visualizing complex traffic environments and historical trajectory information of traffic participants into image prompts— Transportation Context Map (TC-Map), accompanied by corresponding text prompts. Through this approach, we obtained rich traffic context information from the LLM. By integrating this information into the motion prediction model, we demonstrate that such context can enhance the accuracy of motion predictions. Furthermore, considering the cost associated with LLMs, we propose a cost-effective deployment strategy: enhancing the accuracy of motion prediction tasks at scale with 0.7% LLM-augmented datasets. Our research offers valuable insights into enhancing the understanding of traffic scenes of LLMs and the motion prediction performance of autonomous driving. The source code is available at https://github.com/AIR-DISCOVER/LLM-Augmented-MTR and https://aistudio.baidu.com/projectdetail/7809548.

Recommended citation: X Zheng, L Wu, Z Yan, Y Tang, H Zhao, C Zhong, B Chen, J Gong (2024). Large language models powered context-aware motion prediction in autonomous driving. 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems ....

Driving style alignment for llm-powered driver agent Permalink

Published in 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems ..., 2024

Authors: R Yang, X Zhang, A Fernandez-Laaksonen, X Ding, Jiangtao Gong

Abstract: Recently, LLM-powered driver agents have demonstrated considerable potential in the field of autonomous driving, showcasing human-like reasoning and decision-making abilities. However, current research on aligning driver agent behaviors with human driving styles remains limited, partly due to the scarcity of high-quality natural language data from human driving behaviors. To address this research gap, we propose a multi-alignment framework designed to align driver agents with human driving styles through demonstrations and feedback. Notably, we construct a natural language dataset of human driver behaviors through naturalistic driving experiments and post-driving interviews, offering high-quality human demonstrations for LLM alignment. The framework’s effectiveness is validated through simulation experiments in the CARLA urban traffic simulator and further corroborated by human evaluations. Our research offers valuable insights into designing driving agents with diverse driving styles. The implementation of the framework 1 and details of the dataset 2 can be found at the link.

Recommended citation: R Yang, X Zhang, A Fernandez-Laaksonen, X Ding, J Gong (2024). Driving style alignment for llm-powered driver agent. 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems ....

Evaluation of Pedestrian Safety in a High-Fidelity Simulation Environment Framework Permalink

Published in 2024 IEEE 27th International Conference on Intelligent Transportation ..., 2024

Authors: L Ma, L Chen, Y Zhang, M Chu, W Jiang, J Shen, C Li, Y Shi, N Luo, …, Jiangtao Gong

Abstract: Pedestrians' safety is a crucial factor in assessing autonomous driving scenarios. However, pedestrian safety evaluation is rarely considered by existing autonomous driving simulation platforms. This paper proposes a pedestrian safety evaluation method for autonomous driving, in which not only the collision events but also the conflict events together with the characteristics of pedestrians are fully considered. Moreover, to apply the pedestrian safety evaluation system, we construct a high-fidelity simulation framework embedded with pedestrian safety-critical characteristics. We demonstrate our simulation framework and pedestrian safety evaluation with a comparative experiment with two kinds of autonomous driving perception algorithms—single-vehicle perception and vehicle-to-infrastructure (V2I) cooperative perception. The results show that our framework can evaluate different autonomous driving algorithms with detailed and quantitative pedestrian safety indexes. To this end, the proposed simulation method and framework can be used to access different autonomous driving algorithms and evaluate pedestrians' safety performance in future autonomous driving simulations, which can inspire more pedestrian-friendly autonomous driving algorithms.

Recommended citation: L Ma, L Chen, Y Zhang, M Chu, W Jiang, J Shen, C Li, Y Shi, N Luo,... (2024). Evaluation of Pedestrian Safety in a High-Fidelity Simulation Environment Framework. 2024 IEEE 27th International Conference on Intelligent Transportation ....

Understanding Embodied Reference with Touch-Line Transformer. Permalink

Published in ICLR, 2023

Authors: Y Li, X Chen, H Zhao, Jiangtao Gong, G Zhou, F Rossano, Y Zhu

Abstract: We study embodied reference understanding, the task of locating referents using embodied gestural signals and language references. Human studies have revealed that objects referred to or pointed to do not lie on the elbow-wrist line, a common misconception; instead, they lie on the so-called virtual touch line. However, existing human pose representations fail to incorporate the virtual touch line. To tackle this problem, we devise the touch-line transformer: It takes as input tokenized visual and textual features and simultaneously predicts the referent's bounding box and a touch-line vector. Leveraging this touch-line prior, we further devise a geometric consistency loss that encourages the co-linearity between referents and touch lines. Using the touch-line as gestural information improves model performances significantly. Experiments on the YouRefIt dataset show our method achieves a +25.0% accuracy improvement under the 0.75 IoU criterion, closing 63.6% of the gap between model and human performances. Furthermore, we computationally verify prior human studies by showing that computational models more accurately locate referents when using the virtual touch line than when using the elbow-wrist line.

Recommended citation: Y Li, X Chen, H Zhao, J Gong, G Zhou, F Rossano, Y Zhu (2023). Understanding Embodied Reference with Touch-Line Transformer.. ICLR.

Touch-and-heal: Data-driven affective computing in tactile interaction with robotic dog Permalink

Published in Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous ..., 2023

Authors: S Guo, L Zhan, Y Cao, C Zheng, G Zhou, Jiangtao Gong

Recommended citation: S Guo, L Zhan, Y Cao, C Zheng, G Zhou, J Gong (2023). Touch-and-heal: Data-driven affective computing in tactile interaction with robotic dog. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous ....

Int2: Interactive trajectory prediction at intersections Permalink

Published in Proceedings of the IEEE/CVF International Conference on Computer Vision ..., 2023

Authors: Z Yan, P Li, Z Fu, S Xu, Y Shi, X Chen, Y Zheng, Y Li, T Liu, C Li, N Luo, …, Jiangtao Gong

Abstract: Motion forecasting is an important component in autonomous driving systems. One of the most challenging problems in motion forecasting is interactive trajectory prediction, whose goal is to jointly forecasts the future trajectories of interacting agents. To this end, we present a large-scale interactive trajectory prediction dataset named INT2 for INTeractive trajectory prediction at INTersections. INT2 includes 612,000 scenes, each lasting 1 minute, containing up to 10,200 hours of data. The agent trajectories are auto-labeled by a high-performance offline temporal detection and fusion algorithm, whose quality is further inspected by human judges. Vectorized semantic maps and traffic light information are also included in INT2. Additionally, the dataset poses an interesting domain mismatch challenge. For each intersection, we treat rush-hour and non-rush-hour segments as different domains. We benchmark the best open-sourced interactive trajectory prediction method on INT2 and Waymo Open Motion, under in-domain and cross-domain settings. The dataset, code and models are publicly available at https://github.com/AIRDISCOVER/INT2.

Recommended citation: Z Yan, P Li, Z Fu, S Xu, Y Shi, X Chen, Y Zheng, Y Li, T Liu, C Li, N Luo,... (2023). Int2: Interactive trajectory prediction at intersections. Proceedings of the IEEE/CVF International Conference on Computer Vision ....

Enable natural tactile interaction for robot dog based on large-format distributed flexible pressure sensors Permalink

Published in arXiv preprint arXiv:2303.07595, 2023

Authors: L Zhan, Y Cao, Q Chen, H Guo, J Gao, Y Luo, S Guo, G Zhou, Jiangtao Gong

Abstract: Touch is an important channel for human-robot interaction, while it is challenging for robots to recognize human touch accurately and make appropriate responses. In this paper, we design and implement a set of large-format distributed flexible pressure sensors on a robot dog to enable natural human-robot tactile interaction. Through a heuristic study, we sorted out 81 tactile gestures commonly used when humans interact with real dogs and 44 dog reactions. A gesture classification algorithm based on ResNet is proposed to recognize these 81 human gestures, and the classification accuracy reaches 98.7%. In addition, an action prediction algorithm based on Transformer is proposed to predict dog actions from human gestures, reaching a 1-gram BLEU score of 0.87. Finally, we compare the tactile interaction with the voice interaction during a freedom human-robot-dog interactive playing study. The results show that tactile interaction plays a more significant role in alleviating user anxiety, stimulating user excitement and improving the acceptability of robot dogs.

Recommended citation: L Zhan, Y Cao, Q Chen, H Guo, J Gao, Y Luo, S Guo, G Zhou, J Gong (2023). Enable natural tactile interaction for robot dog based on large-format distributed flexible pressure sensors. arXiv preprint arXiv:2303.07595.

M2Sim: A Long-Term Interactive Driving Simulator Permalink

Published in CAAI International Conference on Artificial Intelligence, 172-176, 2023

Authors: Z Han, Z Yan, Y Li, P Li, Y Shi, N Luo, X Gao, Y Shi, P Huang, Jiangtao Gong, …

Recommended citation: Z Han, Z Yan, Y Li, P Li, Y Shi, N Luo, X Gao, Y Shi, P Huang, J Gong,... (2023). M2Sim: A Long-Term Interactive Driving Simulator. CAAI International Conference on Artificial Intelligence, 172-176.

Long-Term Interactive Driving Simulation: MPC to the Rescue Permalink

Published in CAAI International Conference on Artificial Intelligence, 177-188, 2023

Authors: Z Han, Z Yan, Y Li, P Li, Y Shi, N Luo, X Gao, Y Shi, P Huang, Jiangtao Gong, …

Recommended citation: Z Han, Z Yan, Y Li, P Li, Y Shi, N Luo, X Gao, Y Shi, P Huang, J Gong,... (2023). Long-Term Interactive Driving Simulation: MPC to the Rescue. CAAI International Conference on Artificial Intelligence, 177-188.

Planning assembly sequence with graph transformer Permalink

Published in arXiv preprint arXiv:2210.05236, 2022

Authors: L Ma, Jiangtao Gong, H Xu, H Chen, H Zhao, W Huang, G Zhou

Abstract: Assembly sequence planning (ASP) is the essential process for modern manufacturing, proven to be NP-complete thus its effective and efficient solution has been a challenge for researchers in the field. In this paper, we present a graph-transformer based framework for the ASP problem which is trained and demonstrated on a self-collected ASP database. The ASP database contains a self-collected set of LEGO models. The LEGO model is abstracted to a heterogeneous graph structure after a thorough analysis of the original structure and feature extraction. The ground truth assembly sequence is first generated by brute-force search and then adjusted manually to in line with human rational habits. Based on this self-collected ASP dataset, we propose a heterogeneous graph-transformer framework to learn the latent rules for assembly planning. We evaluated the proposed framework in a series of experiment. The results show that the similarity of the predicted and ground truth sequences can reach 0.44, a medium correlation measured by Kendall's $τ$. Meanwhile, we compared the different effects of node features and edge features and generated a feasible and reasonable assembly sequence as a benchmark for further research. Our data set and code is available on https://github.com/AIR-DISCOVER/ICRA\_ASP.

Recommended citation: L Ma, J Gong, H Xu, H Chen, H Zhao, W Huang, G Zhou (2022). Planning assembly sequence with graph transformer. arXiv preprint arXiv:2210.05236.

Learning with yourself: a tangible twin robot system to promote STEM education Permalink

Published in 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems ..., 2022

Authors: J Gao, Jiangtao Gong, G Zhou, H Guo, T Qi

Abstract: This paper presents a customized programmable robotic system, TanTwin (Tangible Twin), designed to promote STEM education for K-12 children. Firstly, TanTwin is implemented based on a wheel-robot with standard LEGO bricks. With several deep neural networks, a child can convert a captured portrait of himself/herself into standard LEGO bricks, therefore he/she can build a tangible twin robot of him-selflherself automatically. Besides, to adapt to the customized appearance, the corresponding visual element and content of the robotic system were also changed by a rule-based adaption algorithm. To demonstrate the effectiveness of TanTwin and to investigate whether tangible twin robots could contribute to children's learning, we conducted a controlled experimental study to compare learning with a TanTwin and with a standard robot system through measuring students' cognitive learning outcomes. The pre-/post- knowledge test results indicated that learning with a tangible twin robot leads to significantly better learning outcomes. Given the results, we validate our system and customization technology can promote STEM education.

Recommended citation: J Gao, J Gong, G Zhou, H Guo, T Qi (2022). Learning with yourself: a tangible twin robot system to promote STEM education. 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems ....

Olap visual analytics on large software call graphs with hierarchical chordmap Permalink

Published in 2015 IEEE International Conference on Data Mining Workshop (ICDMW), 675-679, 2015

Authors: L Wang, Jiangtao Gong, L Shi

Abstract: The performance of call graph analytics is critical to the development and management of large-scale software systems. Classical call graph visualizers either reuse a plain node-link graph metaphor that does not scale, or adopt the specific aggregation technique like the PivotGraph. In this paper, we first generalize the OLAP analysis framework from multidimensional data to multivariate graphs. Then we introduce Hierarchical ChordMap, a new visualization design that tightly couples with OLAP operations through context-preserving user interactions. Controlled user study demonstrates significant improvements of our design from the basic call graph visualizer. Implications are summarized on when and how the ChordMap design can outperform a stable PivotGraph tool in call graph analysis.

Recommended citation: L Wang, J Gong, L Shi (2015). Olap visual analytics on large software call graphs with hierarchical chordmap. 2015 IEEE International Conference on Data Mining Workshop (ICDMW), 675-679.

BugMap: a topographic map of bugs Permalink

Published in Proceedings of the 2013 9th Joint Meeting on Foundations of Software ..., 2013

Authors: Jiangtao Gong, H Zhang

Abstract: A large and complex software system could contain a large number of bugs. It is desirable for developers to understand how these bugs are distributed across the system, so they could have a better overview of software quality. In this paper, we describe BugMap, a tool we developed for visualizing large-scale bug location information. Taken source code and bug data as the input, BugMap can display bug localizations on a topographic map. By examining the topographic map, developers can understand how the components and files are affected by bugs. We apply this tool to visualize the distribution of Eclipse bugs across components/files. The results show that our tool is effective for understanding the overall quality status of a large-scale system and for identifying the problematic areas of the system.

Recommended citation: J Gong, H Zhang (2013). BugMap: a topographic map of bugs. Proceedings of the 2013 9th Joint Meeting on Foundations of Software ....

Education & Learning教育与学习14

How GenAI Mentor Configurations Shape Early Collaborative Dynamics: A Classroom Comparison of Individual and Shared Agents Permalink

Published in arXiv preprint arXiv:2603.12600, 2026

Authors: S Zha, W Liu, F Qin, J Cao, Y Wang, Y Liu, K Zhang, Jiangtao Gong, Y Xu

Abstract: Generative artificial intelligence (GenAI) is increasingly embedded in computer-supported collaborative learning (CSCL), yet little empirical research has unpacked how different configurations of AI participation reshape collaborative processes. This study investigates how GenAI configuration shapes collaborative regulation in authentic classroom settings. Two eighth-grade classes engaged in small-group creative problem-solving under two conditions: a shared-AI configuration, in which each group interacted with a single AI mentor, and an individual-AI configuration, in which each student accessed a personal AI instance. Using multi-layer discourse coding combined with lag sequential analysis (LSA) and ordered network analysis (ONA), we examined interaction distribution, AI-student coupling, shared regulation processes, and teacher orchestration. Results reveal distinct regulatory dynamics across configurations. Shared AI access promoted convergence-oriented collaboration, with stronger alignment of shared regulatory states and more coordinated group-level reasoning. In contrast, individual AI access distributed support across learners, producing more exploratory and evaluative cycles but also more fragmented interaction patterns, accompanied by increased teacher intervention to manage divergence. These findings suggest that AI configuration functions as a structural design variable that reorganizes the regulatory ecology of classroom collaboration.

Recommended citation: S Zha, W Liu, F Qin, J Cao, Y Wang, Y Liu, K Zhang, J Gong, Y Xu (2026). How GenAI Mentor Configurations Shape Early Collaborative Dynamics: A Classroom Comparison of Individual and Shared Agents. arXiv preprint arXiv:2603.12600.

Mentigo: An intelligent agent for mentoring students in the creative problem solving process Permalink

Published in Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems ..., 2025

Authors: S Zha, Y Liu, C Zheng, J Xu, F Yu, Jiangtao Gong, Y Xu

Abstract: Creative Problem-Solving (CPS) promotes creative and critical thinking while enhancing real-world problem-solving skills, making it essential for middle school education.However, providing personalized mentorship in CPS projects at scale is challenging due to resource constraints and diverse student needs.To address this, we developed Mentigo, an AI-driven mentor agent designed to guide middle school students through the CPS process.Using a dataset of real classroom interactions, we encoded CPS task stages, adaptive guidance strategies, and personalized feedback mechanisms to inform Mentigo's dynamic mentoring framework powered by large language models (LLMs).A comparative experiment with 12 students and evaluations from five expert educators demonstrated improved student engagement, creativity, and task performance.Our findings highlight design implications for using LLM-based AI mentors to enhance CPS learning in educational environments.

Recommended citation: S Zha, Y Liu, C Zheng, J Xu, F Yu, J Gong, Y Xu (2025). Mentigo: An intelligent agent for mentoring students in the creative problem solving process. Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems ....

Designing LLM-simulated Immersive Spaces to Enhance Autistic Children’s Social Affordances Understanding in Traffic Settings Permalink

Published in Proceedings of the 30th International Conference on Intelligent User ..., 2025

Authors: Y Cao, Y He, Y Chen, M Chen, S You, Y Qiu, M Liu, C Luo, C Zheng, …, Jiangtao Gong

Abstract: One of the key challenges faced by autistic children is understanding social affordances in complex environments, which further impacts their ability to respond appropriately to social signals. In traffic scenarios, this impairment can even lead to safety concerns. In this paper, we introduce an LLM-simulated immersive projection environment designed to improve this ability in autistic children while ensuring their safety. We first propose 17 design considerations across four major categories, derived from a comprehensive review of previous research. Next, we developed a system called AIroad, which leverages LLMs to simulate drivers with varying social intents, expressed through explicit multimodal social signals. AIroad helps autistic children bridge the gap in recognizing the intentions behind behaviors and learning appropriate responses through various stimuli. A user study involving 14 participants demonstrated that this technology effectively engages autistic children and leads to significant improvements in their comprehension of social affordances in traffic scenarios. Additionally, parents reported high perceived usability of the system. These findings highlight the potential of combining LLM technology with immersive environments for the functional rehabilitation of autistic children in the future.

Recommended citation: Y Cao, Y He, Y Chen, M Chen, S You, Y Qiu, M Liu, C Luo, C Zheng,... (2025). Designing LLM-simulated Immersive Spaces to Enhance Autistic Children's Social Affordances Understanding in Traffic Settings. Proceedings of the 30th International Conference on Intelligent User ....

Designing child-centric AI learning environments: Insights from an LLM-powered creative project-based learning study Permalink

Published in International Journal of Human-Computer Studies, 103602, 2025

Authors: S Zha, Y Qiao, Q Hu, Z Li, Jiangtao Gong, Y Xu

Abstract: Project-based learning (PBL) is an instructional method that is very helpful in nurturing students' creativity, but it requires significant time and energy from both students and teachers. Large language models (LLMs) have been proven to assist in creative tasks, yet much controversy exists regarding their role in fostering creativity. This paper explores the potential of LLMs in PBL settings, with a special focus on fostering creativity. We began with an exploratory study involving 12 middle school students and identified five design considerations for LLM applications in PBL. Building on this, we developed an LLM-empowered, 48-hour PBL program and conducted an instructional experiment with 31 middle school students. Our results indicated that LLMs can enhance every stage of PBL. Additionally, we also discovered ambivalent perspectives among students and mentors toward LLM usage. Furthermore, we explored the challenge and design implications of integrating LLMs into PBL and reflected on the program. By bridging AI advancements into educational practice, our work aims to inspire further discourse and investigation into harnessing AI's potential in child-centric educational settings.

Recommended citation: S Zha, Y Qiao, Q Hu, Z Li, J Gong, Y Xu (2025). Designing child-centric AI learning environments: Insights from an LLM-powered creative project-based learning study. International Journal of Human-Computer Studies, 103602.

COLP: scaffolding children’s online long-term collaborative learning Permalink

Published in International Journal of Human-Computer Interaction 41 (24), 15453-15475, 2025

Authors: S Zha, Y Tang, Jiangtao Gong, Y Xu

Abstract: Online collaborative learning and working are important for everyone including children. However, children still face a lot of difficulties communicating and working together while online, which keeps them from engaging in long-term project-based teamwork. We aim to investigate online long-term collaborative learning opportunities to address this gap. We design COLP, an online, 16-week, project-based learning program, as an educational intervention based on multiple learning theories for primary school students. We conducted this program with 67 primary school students ages 8-13, across more than five provinces of China. We found that this program could engage more than one-third of children in teamwork after long-term study. Furthermore, we interview children and their parents to help us understand the communication channel, benefits, and challenges of this program. Interestingly, we discovered that parents play multiple roles in their children's collaborative learning, particularly modeling and guiding the children's collaborative skills. Given the lack of programs designed for children's long-term online collaboration, this study may inspire intervention design in computer-supported collaborative learning communities.

Recommended citation: S Zha, Y Tang, J Gong, Y Xu (2025). COLP: scaffolding children's online long-term collaborative learning. International Journal of Human-Computer Interaction 41 (24), 15453-15475.

Guiding Multiple Remote Users in Physical Tasks with Language-driven Robotic Telepresence Permalink

Published in Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous ..., 2025

Authors: R Li, J Guo, X Zhang, X Zhang, Z Li, J Li, Jiangtao Gong

Abstract: Remote assistance through robotic telepresence could involve both control and memory challenges, particularly in one expert to multiple workers situation. In this work, we proposed a novelty language-driven interface to facilitate remote collaboration through telepresence robots. Through operations and maintenance expert interviews and a scenario simulation study, we identified key pain points in executing one-expert-multiple-workers remote guidance using the telepresence robot and proposed two design goals, which together consist of five sub-design goals with corresponding features. These features were integrated into a standard telepresence robot, resulting in the development of a Collaborative LLM-based Embodied Assistant Robot, named CLEAR Robot. A controlled experiment simulating a remote assembly task of one to two demonstrated that, compared to the standard telepresence robot, CLEAR Robot significantly improved efficiency, reduced cognitive load, facilitated more balanced collaboration, and improved the user experience. We also discuss the impact of language-driven implicit interactions in multi-user collaboration and provide insights for designing robot systems that support one-expert-multiple-workers remote guidance in the future.

Recommended citation: R Li, J Guo, X Zhang, X Zhang, Z Li, J Li, J Gong (2025). Guiding Multiple Remote Users in Physical Tasks with Language-driven Robotic Telepresence. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous ....

Teleaware robot: Designing awareness-augmented telepresence robot for remote collaborative locomotion Permalink

Published in Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous ..., 2024

Authors: R Li, Y Zhu, M Liu, Y Zeng, S Zhuang, J Fu, Y Lu, G Zhou, C Liu, Jiangtao Gong

Abstract: Telepresence robots can be used to support users to navigate an environment remotely and share the visiting experience with their social partners. Although such systems allow users to see and hear the remote environment and communicate with their partners via live video feed, this does not provide enough awareness of the environment and their remote partner's activities. In this paper, we introduce an awareness framework for collaborative locomotion in scenarios of onsite and remote users visiting a place together. From an observational study of small groups of people visiting exhibitions, we derived four design goals for enhancing the environmental and social awareness between social partners, and developed a set of awareness-enhancing techniques to add to a standard telepresence robot - named TeleAware robot. Through a controlled experiment simulating a guided exhibition visiting task, TeleAware robot showed the ability to lower the workload, facilitate closer social proximity, and improve mutual awareness and social presence compared with the standard one. We discuss the impact of mobility and roles of local and remote users, and provide insights for the future design of awareness-enhancing telepresence robot systems that facilitate collaborative locomotion.

Recommended citation: R Li, Y Zhu, M Liu, Y Zeng, S Zhuang, J Fu, Y Lu, G Zhou, C Liu, J Gong (2024). Teleaware robot: Designing awareness-augmented telepresence robot for remote collaborative locomotion. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous ....

MR. Brick: designing a remote mixed-reality educational game system for promoting children’s social & collaborative skills Permalink

Published in Proceedings of the 2023 CHI conference on human factors in computing systems ..., 2023

Authors: Y Wu, S You, Z Guo, X Li, G Zhou, Jiangtao Gong

Abstract: Children are one of the groups most influenced by COVID-19-related social distancing, and a lack of contact with peers can limit their opportunities to develop social and collaborative skills. However, remote socialization and collaboration as an alternative approach is still a great challenge for children. This paper presents MR.Brick, a Mixed Reality (MR) educational game system that helps children adapt to remote collaboration. A controlled experimental study involving 24 children aged six to ten was conducted to compare MR.Brick with the traditional video game by measuring their social and collaborative skills and analyzing their multi-modal playing behaviours. The results showed that MR.Brick was more conducive to children’s remote collaboration experience than the traditional video game. Given the lack of training systems designed for children to collaborate remotely, this study may inspire interaction design and educational research in related fields.

Recommended citation: Y Wu, S You, Z Guo, X Li, G Zhou, J Gong (2023). MR. Brick: designing a remote mixed-reality educational game system for promoting children's social & collaborative skills. Proceedings of the 2023 CHI conference on human factors in computing systems ....

ChordAR: An educational AR game design for children’s music theory learning Permalink

Published in Wireless communications and mobile computing 2022 (1), 5268586, 2022

Authors: Y Lu, X Wang, Jiangtao Gong, Y Liang

Abstract: Augmented reality (AR) technology, with its unique immersive interactive experience of virtual reality integration, can show children abstract concepts and theories in an intuitive way. In this paper, we improve the high‐precision position and low delay interaction of AR recognition area through edge calculation, so as to enhance the immersion of children’s chord knowledge learning. Meanwhile, multisource fusion perception of audio and touch is used, and the AI algorithm model is used to monitor physical events (cloud AI training and edge execution) and effectively realize intelligent services such as chord combination, melody perception, and control. Chord knowledge in music theory is presented to children through multichannels (kinesthetic, visual, auditory, etc.) situationally, to improve children’s learning efficiency and experience. We invited 12 children to participate in user experiments to test usability of ChordAR. The test results showed that children can master the method of playing the game through simple learning, ultimately correctly imitating C, G, and E chords, and experience full immersion in the game. Finally, we discuss the positive effects of ChordAR on children’s learning and creativity and make suggestions for future AR games.

Recommended citation: Y Lu, X Wang, J Gong, Y Liang (2022). ChordAR: An educational AR game design for children's music theory learning. Wireless communications and mobile computing 2022 (1), 5268586.

Holoboard: A large-format immersive teaching board based on pseudo holographics Permalink

Published in The 34th Annual ACM Symposium on User Interface Software and Technology, 441-456, 2021

Authors: Jiangtao Gong, T Han, S Guo, J Li, S Zha, L Zhang, F Tian, Q Wang, Y Rui

Abstract: In this paper, we present HoloBoard, an interactive large-format pseduo-holographic display system for lecture based classes. With its unique properties of immersive visual display and transparent screen, we designed and implemented a rich set of novel interaction techniques like immersive presentation, role-play, and lecturing behind the scene that are potentially valuable for lecturing in class. We conducted a controlled experimental study to compare a HoloBoard class with a normal class through measuring students’ learning outcomes and three dimensions of engagement (i.e., behavioral, emotional, and cognitive engagement). We used pre-/post- knowledge tests and multimodal learning analytics to measure students’ learning outcomes and learning experiences. Results indicated that the lecture-based class utilizing HoloBoard lead to slightly better learning outcomes and a significantly higher level of student engagement. Given the results, we discussed the impact of HoloBoard as an immersive media in the classroom setting and suggest several design implications for deploying HoloBoard in immersive teaching practices.

Recommended citation: J Gong, T Han, S Guo, J Li, S Zha, L Zhang, F Tian, Q Wang, Y Rui (2021). Holoboard: A large-format immersive teaching board based on pseudo holographics. The 34th Annual ACM Symposium on User Interface Software and Technology, 441-456.

HoloBoard: an Immersive Teaching Board System Permalink

Published in Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing ..., 2021

Authors: Jiangtao Gong, X Ji, S Zha, T Han, Q Ding, J Li, L Zhang, Q Wang

Abstract: We present HoloBoard, a next generation teaching board based on large-format, semi-transparent, interactive display technology that supports immersive teaching and learning. It can present immersive media curriculum materials including large-format videos, interactive 3D objects, AR effects, etc. Based on a multi-stage mixed-methods teacher-centered design thinking approach and a series of user studies with 6-9 years old primary students, the concept design of HoloBoard was defined and refined. The results showed that HoloBoard can be deployed in the wild and bring more engagement and interactivity in the primary school classroom, promoting students’ spontaneous exploration and active learning.

Recommended citation: J Gong, X Ji, S Zha, T Han, Q Ding, J Li, L Zhang, Q Wang (2021). HoloBoard: an Immersive Teaching Board System. Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing ....

PaperLego: Component-based papercraft designing tool for children Permalink

Published in SIGGRAPH Asia 2014 Designing Tools For Crafting Interactive Artifacts, 1-4, 2014

Authors: Jiangtao Gong, J Wang, Y Xu

Abstract: In this paper, we propose a new papercraft designing tool for children, PaperLego, which can make arbitrary 3D Objects to 2D paper folding patterns for children. PaperLego use some simple 3D primitives such as frustums or hexahedrons as components, and when glued together, they form 3D object that closely resembles the original model. PaperLego can also support children to create based on the original model and the paper model. The main problem we address in this paper is what kind of paper folding patterns is suitable for children and how we produce this kind of patterns by computer algorithms. Our user study results show that young children at the age of 7 can successfully use the paper folding patterns generated by our tool to build a wide variety of complex 3D models, and the folding process is intuitive and fun.

Recommended citation: J Gong, J Wang, Y Xu (2014). PaperLego: Component-based papercraft designing tool for children. SIGGRAPH Asia 2014 Designing Tools For Crafting Interactive Artifacts, 1-4.

Traffic & Autonomous Driving交通与自动驾驶2

CoIRL-AD: Collaborative-Competitive Imitation-Reinforcement Learning in Latent World Models for Autonomous Driving Permalink

Published in arXiv preprint arXiv:2510.12560, 2025

Authors: X Zheng, Z Yang, Y Chen, Y Peng, Y Tang, G Liu, B Chen, Jiangtao Gong

Abstract: End-to-end autonomous driving models trained with imitation learning (IL) often generalize poorly, particularly in long-tail scenarios where expert demonstrations are sparse. Reinforcement learning (RL) can provide complementary task-level supervision, but applying RL to real-world autonomous driving is challenging in offline settings without interactive simulators, where datasets are dominated by expert actions and provide limited behavioral diversity. We propose CoIRL-AD, a competitive dual-policy framework that integrates IL and RL under a unified offline training regime. CoIRL-AD decouples imitation and reward optimization into separate actors to alleviate objective conflicts, uses imagined future rollouts for long-horizon reward estimation, and introduces a competition mechanism that selectively transfers beneficial behaviors while keeping RL anchored to expert-like driving. Experiments on the nuScenes benchmark show that CoIRL-AD consistently improves robustness over strong IL-based baselines, with especially large gains in cross-city generalization and long-tail scenarios. Code is available at: https://github.com/SEU-zxj/CoIRL-AD.

Recommended citation: X Zheng, Z Yang, Y Chen, Y Peng, Y Tang, G Liu, B Chen, J Gong (2025). CoIRL-AD: Collaborative-Competitive Imitation-Reinforcement Learning in Latent World Models for Autonomous Driving. arXiv preprint arXiv:2510.12560.

A High Fidelity Simulation Framework for Potential Safety Benefits Estimation of Cooperative Pedestrian Perception Permalink

Published in arXiv e-prints, arXiv: 2210.08731, 2022

Authors: L Chen, Y Zhang, W Jiang, Jiangtao Gong, J Shen, M Chu, C Li, Y Pan, Y Shi, …

Recommended citation: L Chen, Y Zhang, W Jiang, J Gong, J Shen, M Chu, C Li, Y Pan, Y Shi,... (2022). A High Fidelity Simulation Framework for Potential Safety Benefits Estimation of Cooperative Pedestrian Perception. arXiv e-prints, arXiv: 2210.08731.

Accessibility无障碍9

AI system facilitates people with blindness and low vision in interpreting and experiencing unfamiliar environments Permalink

Published in npj Artificial Intelligence 1 (1), 7, 2025

Authors: H Lin, Jiangtao Gong, Y Wang, J Zhang, B Bai, Y Zhang, L Wang, C Wei, Y Cao, …

Abstract: Engaging with nature significantly enhances well-being, yet millions of individuals with blindness and low vision (BLV) are often excluded from these benefits due to constrained environmental perception. Here, we introduce VIPTour, an AI-driven system powered by the FocusFormer algorithm, which transforms complex scenes into structured, personalized graphs using tailored attention mechanisms and a BLV-in-the-Loop Adapter. Through intuitive, user-centered interaction, VIPTour facilitates active exploration and in-depth comprehension during dynamic sightseeing, while enabling accurate, long-lasting recollection and effective communication among BLV individuals post-journey. Extensive experiments demonstrate that VIPTour significantly enhances positive emotions and memory retention, with a 67.9% increase in positive emotional response, a 94.7% rise in arousal, a 772.73% improvement in cognitive mapping accuracy, and a 200% enhancement in long-term memory accuracy. These results underscore VIPTour’s ability to deliver an unparalleled, enjoyable, and memorable experience, promising profound benefits for the BLV community.

Recommended citation: H Lin, J Gong, Y Wang, J Zhang, B Bai, Y Zhang, L Wang, C Wei, Y Cao,... (2025). AI system facilitates people with blindness and low vision in interpreting and experiencing unfamiliar environments. npj Artificial Intelligence 1 (1), 7.

“I am the follower, also the boss”: Exploring Different Levels of Autonomy and Machine Forms of Guiding Robots for the Visually Impaired Permalink

Published in Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems ..., 2023

Authors: Y Zhang, Z Li, H Guo, L Wang, Q Chen, W Jiang, M Fan, G Zhou, Jiangtao Gong

Abstract: Guiding robots, in the form of canes or cars, have recently been explored to assist blind and low vision (BLV) people. Such robots can provide full or partial autonomy when guiding. However, the pros and cons of different forms and autonomy for guiding robots remain unknown. We sought to fill this gap. We designed autonomy-switchable guiding robotic cane and car. We conducted a controlled lab-study (N=12) and a field study (N=9) on BLV. Results showed that full autonomy received better walking performance and subjective ratings in the controlled study, whereas participants used more partial autonomy in the natural environment as demanding more control. Besides, the car robot has demonstrated abilities to provide a higher sense of safety and navigation efficiency compared with the cane robot. Our findings offered empirical evidence about how the BLV community perceived different machine forms and autonomy, which can inform the design of assistive robots.

Recommended citation: Y Zhang, Z Li, H Guo, L Wang, Q Chen, W Jiang, M Fan, G Zhou, J Gong (2023). " I am the follower, also the boss": Exploring Different Levels of Autonomy and Machine Forms of Guiding Robots for the Visually Impaired. Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems ....

HeliCoach: An Adaptive Orientation and Mobility Training System in a Drone-Based 3D Audio Space Permalink

Published in Extended Abstracts of the 2020 CHI Conference on Human Factors in Computing ..., 2020

Authors: Jiangtao Gong, Q Ding, Y Zhang, L Zhang, P Xu, M Liu, Y Xu, Q Wang

Abstract: We present HeliCoach, an adaptive O&M (orientation and mobility) training system to help visually impaired people to gain audio-orientation ability efficiently. HeliCoach is a mastery-based adaptive learning system consists of three components: A drone-based physically simulated 3D audio space, a UWB (Ultra-Wide Band)-supported assessment system, an intelligent belt with haptic feedback as training scaffolds. We evaluated the acceptability and efficiency of HeliCoach on both blindfolded sighted participants and visually impaired participants in an audio-orientation task. The results showed that HeliCoach was efficient and interesting. Thus, We believe HeliCoach can be applied to many simulated 3D audio space related task to train people more efficiently.

Recommended citation: J Gong, Q Ding, Y Zhang, L Zhang, P Xu, M Liu, Y Xu, Q Wang (2020). HeliCoach: An Adaptive Orientation and Mobility Training System in a Drone-Based 3D Audio Space. Extended Abstracts of the 2020 CHI Conference on Human Factors in Computing ....

HeliCoach: An Adaptive Multimodal Orientation and Mobility Training System in a Drone-Based Simulated 3D Audio Space Permalink

Published in Journal of Computer-Aided Design & Computer Graphics 32 (7), 1129-1136, 2020

Authors: Jiangtao Gong, Q Ding, P Xu, Y Zhang, L Zhang, Q Wang

Recommended citation: J Gong, Q Ding, P Xu, Y Zhang, L Zhang, Q Wang (2020). HeliCoach: An Adaptive Multimodal Orientation and Mobility Training System in a Drone-Based Simulated 3D Audio Space. Journal of Computer-Aided Design & Computer Graphics 32 (7), 1129-1136.

HeliCoach: 基于无人机模拟 3D 音频空间的多通道自适应定向行走训练系统的设计研究 Permalink

Published in 计算机辅助设计与图形学学报 32 (7), 1129-1136, 2020

Authors: 龚江涛,丁琦城,徐鹏辉,张宇,张柳新,王茜莺

Recommended citation: 龚江涛,丁琦城,徐鹏辉,张宇,张柳新,王茜莺 (2020). HeliCoach: 基于无人机模拟 3D 音频空间的多通道自适应定向行走训练系统的设计研究. 计算机辅助设计与图形学学报 32 (7), 1129-1136.

HeliCoach: A drone-based system supporting orientation and mobility training for the visually impaired Permalink

Published in Extended Abstracts of the 2019 CHI Conference on Human Factors in Computing ..., 2019

Authors: Q Ding, Jiangtao Gong, J Sun, P Xu, L Zhang, Q Wang, Y Zhang

Abstract: Orientation and mobility (O&M) training is essential for improving the independence and wellbeing of the visually impaired. However, the shortage of qualified trainers and the unengaging training contents limit the number of O&M training recipients. In this paper, we propose a drone-based intelligent tutor system - HeliCoach - to provide cost-effective and personalized O&M training. We first elaborate on the system design and potential scenarios of HeliCoach use. We then demonstrate the effectiveness of this concept using a preliminary user study. Finally, we discuss the implication and challenges of this system.

Recommended citation: Q Ding, J Gong, J Sun, P Xu, L Zhang, Q Wang, Y Zhang (2019). HeliCoach: A drone-based system supporting orientation and mobility training for the visually impaired. Extended Abstracts of the 2019 CHI Conference on Human Factors in Computing ....

盲人触觉图像显示器的交互体验研究 Permalink

Published in 计算机辅助设计与图形学学报 28 (9), 1571-1576, 2016

Authors: 丁琦城,龚江涛,史元春,徐迎庆

Recommended citation: 丁琦城,龚江涛,史元春,徐迎庆 (2016). 盲人触觉图像显示器的交互体验研究. 计算机辅助设计与图形学学报 28 (9), 1571-1576.

Well-being身心健康5

Human-Centered AI for Expressive Arts Therapy: Designing for Mental Well-Being and Creativity Permalink

Published in Companion Publication of the 2026 ACM Designing Interactive Systems ..., 2026

Authors: Y Jin, P An, W Cai, J Yang, L Chen, K Moylan, K Doherty, G Doherty, …, Jiangtao Gong

Abstract: Expressive arts therapy supports creative expression through visual art, music, dance, and drama, fostering emotional awareness, regulation, and personal growth. Recent advances in generative AI, particularly multimodal models that produce images, music, text, and soundscapes, create new opportunities to scaffold creative exploration, provide adaptive prompts, and support reflective dialogue in therapeutic contexts. At the same time, integrating AI into expressive arts therapy raises important challenges around agency, authorship, emotional alignment, ethics, data privacy, and preserving therapeutic relationships. This workshop brings together researchers, designers, therapists, and practitioners to explore how human-centered AI can support creativity in art and music therapy while respecting therapeutic values. We focus on insights from AI-enabled therapy systems, design approaches balancing creativity and emotional well-being, and evaluation methods that extend beyond usability and traditional clinical outcomes. Through interdisciplinary discussion, the workshop aims to surface open research questions and design considerations for AI use in expressive arts therapy.

Recommended citation: Y Jin, P An, W Cai, J Yang, L Chen, K Moylan, K Doherty, G Doherty,... (2026). Human-Centered AI for Expressive Arts Therapy: Designing for Mental Well-Being and Creativity. Companion Publication of the 2026 ACM Designing Interactive Systems ....

OYaYa: A Desktop Robot Enabling Multimodal Interaction with Emotions Permalink

Published in Companion Proceedings of the 26th International Conference on Intelligent ..., 2021

Authors: Y Jin, Y Deng, Jiangtao Gong, X Wan, G Gao, Q Wang

Abstract: We demonstrate a desktop robot OYaYa that imitates users’ emotional facial expressions and helps users manage emotions. Multiple equipped sensors in OYaYa enable multimodal interaction; for example, it recognizes users’ emotions from facial expressions and speeches. Besides, a dashboard illustrates how users interact with OYaYa and how their emotions change. We expect that OYaYa allows users to manage their emotions in a fun way.

Recommended citation: Y Jin, Y Deng, J Gong, X Wan, G Gao, Q Wang (2021). OYaYa: A Desktop Robot Enabling Multimodal Interaction with Emotions. Companion Proceedings of the 26th International Conference on Intelligent ....

Cold comfort matters-how channel-wise emotional strategies help in a customer service chatbot Permalink

Published in Extended Abstracts of the 2020 CHI Conference on Human Factors in Computing ..., 2020

Authors: M Liu, Q Ding, Y Zhang, G Zhao, C Hu, Jiangtao Gong, P Xu, Y Zhang, L Zhang, …

Abstract: Smart service chatbots, aiming to provide efficient customer service, have increased rapidly. Currently, few task-oriented chatbots had employed emotion strategies together with problem solving process. Meanwhile, customers can have different preference across channels (e.g., social media and webpages). This paper presented the design for a series of emotional strategies (ES) combined with task-solving, and tested its effectiveness across channels. In the multi-channel service chatbot, Moli, we compared ES with brief apology (BA) and no emotion responses (NER) in a Wizard of Oz study. The results indicated the effectiveness of the emotion strategies borrowed from psychotherapy technologies. Meanwhile, we found customers preferred brief and direct strategies on webpage, whereas rich emotional strategies in all stages were more favorable on social media. With detailed strategy design, the "cold comfort" was found helpful for building human-machine relationships, pacifying emotions, and improving satisfaction and system usability.

Recommended citation: M Liu, Q Ding, Y Zhang, G Zhao, C Hu, J Gong, P Xu, Y Zhang, L Zhang,... (2020). Cold comfort matters-how channel-wise emotional strategies help in a customer service chatbot. Extended Abstracts of the 2020 CHI Conference on Human Factors in Computing ....

Escape from the Dark Jungle: A 3D Audio Game for Emotion Regulation Permalink

Published in International Conference on Virtual, Augmented and Mixed Reality, 57-76, 2018

Authors: Jiangtao Gong, Y Shi, J Wang, D Shi, Y Xu

Recommended citation: J Gong, Y Shi, J Wang, D Shi, Y Xu (2018). Escape from the Dark Jungle: A 3D Audio Game for Emotion Regulation. International Conference on Virtual, Augmented and Mixed Reality, 57-76.

Human-AI Collaboration人机协作11

FreeAskWorld: An Interactive and Closed-Loop Simulator for Human-Centric Embodied AI Permalink

Published in Proceedings of the AAAI Conference on Artificial Intelligence 40 (2), 926-934, 2026

Authors: Y Peng, Y Pan, X He, J Yang, X Yin, H Wang, X Zheng, C Gao, Jiangtao Gong

Abstract: As embodied intelligence emerges as a core frontier in artificial intelligence research, simulation platforms must evolve beyond low-level physical interactions to capture complex, human-centered social behaviors. We introduce FreeAskWorld, an interactive simulation framework that integrates large language models (LLMs) for high-level behavior planning and semantically grounded interaction, informed by theories of intention and social cognition. Our framework supports scalable, realistic human-agent simulations and includes a modular data generation pipeline tailored for diverse embodied tasks.To validate the framework, we extend the classic Vision-and-Language Navigation (VLN) task into a semantically enriched Direction Inquiry setting, wherein agents can actively seek and interpret navigational guidance. We present and publicly release FreeAskWorld, a large-scale benchmark dataset comprising reconstructed environments, six diverse task types, 16 core object categories, 63,429 annotated sample frames, and more than 17 hours of interaction data to support training and evaluation of embodied AI systems. We benchmark VLN models, and human participants under both open-loop and closed-loop settings. Experimental results demonstrate that models fine-tuned on FreeAskWorld outperform their original counterparts, achieving enhanced semantic understanding and interaction competency. These findings underscore the efficacy of socially grounded simulation frameworks in advancing embodied AI systems toward sophisticated high-level planning and more naturalistic human-agent interaction.

Recommended citation: Y Peng, Y Pan, X He, J Yang, X Yin, H Wang, X Zheng, C Gao, J Gong (2026). FreeAskWorld: An Interactive and Closed-Loop Simulator for Human-Centric Embodied AI. Proceedings of the AAAI Conference on Artificial Intelligence 40 (2), 926-934.

Human-Centered AI for Expressive Arts Therapy: Designing for Mental Well-Being and Creativity Permalink

Published in Companion Publication of the 2026 ACM Designing Interactive Systems ..., 2026

Authors: Y Jin, P An, W Cai, J Yang, L Chen, K Moylan, K Doherty, G Doherty, …, Jiangtao Gong

Abstract: Expressive arts therapy supports creative expression through visual art, music, dance, and drama, fostering emotional awareness, regulation, and personal growth. Recent advances in generative AI, particularly multimodal models that produce images, music, text, and soundscapes, create new opportunities to scaffold creative exploration, provide adaptive prompts, and support reflective dialogue in therapeutic contexts. At the same time, integrating AI into expressive arts therapy raises important challenges around agency, authorship, emotional alignment, ethics, data privacy, and preserving therapeutic relationships. This workshop brings together researchers, designers, therapists, and practitioners to explore how human-centered AI can support creativity in art and music therapy while respecting therapeutic values. We focus on insights from AI-enabled therapy systems, design approaches balancing creativity and emotional well-being, and evaluation methods that extend beyond usability and traditional clinical outcomes. Through interdisciplinary discussion, the workshop aims to surface open research questions and design considerations for AI use in expressive arts therapy.

Recommended citation: Y Jin, P An, W Cai, J Yang, L Chen, K Moylan, K Doherty, G Doherty,... (2026). Human-Centered AI for Expressive Arts Therapy: Designing for Mental Well-Being and Creativity. Companion Publication of the 2026 ACM Designing Interactive Systems ....

Human Tool: An MCP-Style Framework for Human-Agent Collaboration Permalink

Published in arXiv preprint arXiv:2602.12953, 2026

Authors: Y Tang, H Peng, B Zhao, H Ding, H Song, T Wang, C Zhong, Jiangtao Gong

Abstract: Human-AI collaboration faces growing challenges as AI systems increasingly outperform humans on complex tasks, while humans remain responsible for orchestration, validation, and decision oversight. To address this imbalance, we introduce Human Tool, an MCP-style interface abstraction, building on recent Model Context Protocol designs, that exposes humans as callable tools within AI-led, proactive workflows. Here, "tool" denotes a coordination abstraction, not a reduction of human authority or responsibility. Building on LLM-based agent architectures, we operationalize Human Tool by modeling human contributions through structured tool schemas of capabilities, information, and authority. These schemas enable agents to dynamically invoke human input based on relative strengths and reintegrate it through efficient, natural interaction protocols. We validate the framework through controlled studies in both decision-making and creative tasks, demonstrating improved task performance, reduced human workload, and more balanced collaboration dynamics compared to baseline systems. Finally, we discuss implications for human-centered AI design, highlighting how MCP-style human tools enable strong AI leadership while amplifying uniquely human strengths.

Recommended citation: Y Tang, H Peng, B Zhao, H Ding, H Song, T Wang, C Zhong, J Gong (2026). Human Tool: An MCP-Style Framework for Human-Agent Collaboration. arXiv preprint arXiv:2602.12953.

Human and algorithmic visual attention in driving tasks Permalink

Published in npj Artificial Intelligence 2 (1), 23, 2026

Authors: C Zheng, P Li, B Jin, S You, KI Chan, YQ Zhang, G Zhou, Jiangtao Gong

Abstract: Amid the heated debate on whether artificial intelligence possesses a human-like capacity for understanding, the compatibility and interaction between human and algorithmic visual attention remain unclear. Here, we address this issue through the lens of spatial and feature-based attention. Using autonomous driving as an epitome of safety-critical domains, we show that human attention in driving tasks can be divided into three phases, each characterized by spatial, feature-based, and mixed visual attention. Comparisons between each phase of human attention and algorithmic attention revealed a complex landscape of human-AI resemblance. For specialized detection and planning algorithms, incorporating semantic-rich, feature-based human attention markedly enhanced performance, suggesting these models lack human-like semantic visual understanding. In contrast, for large-scale Vision-Language Models, the effect was task-dependent, suggesting that while foundation models have bridged the “reasoning gap” through massive pre-training, a “grounding gap” persists in fine-grained visual tasks. Crucially, our findings demonstrate that incorporating human semantic attention offers an effective and economic pathway to compensating for these gaps, enhancing model understanding in safety-critical and grounding-heavy tasks without the need for massive scale.

Recommended citation: C Zheng, P Li, B Jin, S You, KI Chan, YQ Zhang, G Zhou, J Gong (2026). Human and algorithmic visual attention in driving tasks. npj Artificial Intelligence 2 (1), 23.

Designing child-centric AI learning environments: Insights from an LLM-powered creative project-based learning study Permalink

Published in International Journal of Human-Computer Studies, 103602, 2025

Authors: S Zha, Y Qiao, Q Hu, Z Li, Jiangtao Gong, Y Xu

Abstract: Project-based learning (PBL) is an instructional method that is very helpful in nurturing students' creativity, but it requires significant time and energy from both students and teachers. Large language models (LLMs) have been proven to assist in creative tasks, yet much controversy exists regarding their role in fostering creativity. This paper explores the potential of LLMs in PBL settings, with a special focus on fostering creativity. We began with an exploratory study involving 12 middle school students and identified five design considerations for LLM applications in PBL. Building on this, we developed an LLM-empowered, 48-hour PBL program and conducted an instructional experiment with 31 middle school students. Our results indicated that LLMs can enhance every stage of PBL. Additionally, we also discovered ambivalent perspectives among students and mentors toward LLM usage. Furthermore, we explored the challenge and design implications of integrating LLMs into PBL and reflected on the program. By bridging AI advancements into educational practice, our work aims to inspire further discourse and investigation into harnessing AI's potential in child-centric educational settings.

Recommended citation: S Zha, Y Qiao, Q Hu, Z Li, J Gong, Y Xu (2025). Designing child-centric AI learning environments: Insights from an LLM-powered creative project-based learning study. International Journal of Human-Computer Studies, 103602.

Understanding human-AI collaboration in music therapy through co-design with therapists Permalink

Published in Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems ..., 2024

Authors: J Sun, J Yang, G Zhou, Y Jin, Jiangtao Gong

Abstract: The rapid development of musical AI technologies has expanded the creative potential of various musical activities, ranging from music style transformation to music generation. However, little research has investigated how musical AIs can support music therapists, who urgently need new technology support. This study used a mixed method, including semi-structured interviews and a participatory design approach. By collaborating with music therapists, we explored design opportunities for musical AIs in music therapy. We presented the co-design outcomes involving the integration of musical AIs into a music therapy process, which was developed from a theoretical framework rooted in emotion-focused therapy. After that, we concluded the benefits and concerns surrounding music AIs from the perspective of music therapists. Based on our findings, we discussed the opportunities and design implications for applying musical AIs to music therapy. Our work offers valuable insights for developing human-AI collaborative music systems in therapy involving complex procedures and specific requirements.

Recommended citation: J Sun, J Yang, G Zhou, Y Jin, J Gong (2024). Understanding human-AI collaboration in music therapy through co-design with therapists. Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems ....

“It Must Be Gesturing Towards Me”: Gesture-Based Interaction between Autonomous Vehicles and Pedestrians Permalink

Published in Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems ..., 2024

Authors: X Chang, Z Chen, X Dong, Y Cai, T Yan, H Cai, Z Zhou, G Zhou, Jiangtao Gong

Abstract: Interacting with pedestrians understandably and efficiently is one of the toughest challenges faced by autonomous vehicles (AVs) due to the limitations of current algorithms and external human-machine interfaces (eHMIs). In this paper, we design eHMIs based on gestures inspired by the most popular method of interaction between pedestrians and human drivers. Eight common gestures were selected to convey AVs’ yielding or non-yielding intentions at uncontrolled crosswalks from previous literature. Through a VR experiment (N1 = 31) and a following online survey (N2 = 394), we discovered significant differences in the usability of gesture-based eHMIs compared to current eHMIs. Good gesture-based eHMIs increase the efficiency of pedestrian-AV interaction while ensuring safety. Poor gestures, however, cause misinterpretation. The underlying reasons were explored: ambiguity regarding the recipient of the signal and whether the gestures are precise, polite, and familiar to pedestrians. Based on this empirical evidence, we discuss potential opportunities and provide valuable insights into developing comprehensible gesture-based eHMIs in the future to support better interaction between AVs and other road users.

Recommended citation: X Chang, Z Chen, X Dong, Y Cai, T Yan, H Cai, Z Zhou, G Zhou, J Gong (2024). " It Must Be Gesturing Towards Me": Gesture-Based Interaction between Autonomous Vehicles and Pedestrians. Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems ....

“I am the follower, also the boss”: Exploring Different Levels of Autonomy and Machine Forms of Guiding Robots for the Visually Impaired Permalink

Published in Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems ..., 2023

Authors: Y Zhang, Z Li, H Guo, L Wang, Q Chen, W Jiang, M Fan, G Zhou, Jiangtao Gong

Abstract: Guiding robots, in the form of canes or cars, have recently been explored to assist blind and low vision (BLV) people. Such robots can provide full or partial autonomy when guiding. However, the pros and cons of different forms and autonomy for guiding robots remain unknown. We sought to fill this gap. We designed autonomy-switchable guiding robotic cane and car. We conducted a controlled lab-study (N=12) and a field study (N=9) on BLV. Results showed that full autonomy received better walking performance and subjective ratings in the controlled study, whereas participants used more partial autonomy in the natural environment as demanding more control. Besides, the car robot has demonstrated abilities to provide a higher sense of safety and navigation efficiency compared with the cane robot. Our findings offered empirical evidence about how the BLV community perceived different machine forms and autonomy, which can inform the design of assistive robots.

Recommended citation: Y Zhang, Z Li, H Guo, L Wang, Q Chen, W Jiang, M Fan, G Zhou, J Gong (2023). " I am the follower, also the boss": Exploring Different Levels of Autonomy and Machine Forms of Guiding Robots for the Visually Impaired. Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems ....

Work with AI and work for AI: autonomous vehicle safety drivers… lived experiences Permalink

Published in Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems ..., 2023

Authors: M Chu, K Zong, X Shu, Jiangtao Gong, Z Lu, K Guo, X Dai, G Zhou

Abstract: The development of Autonomous Vehicle (AV) has created a novel job, the safety driver, recruited from experienced drivers to supervise and operate AV in numerous driving missions. Safety drivers usually work with non-perfect AV in high-risk real-world traffic environments for road testing tasks. However, this group of workers is under-explored in the HCI community. To fill this gap, we conducted semi-structured interviews with 26 safety drivers. Our results present how safety drivers cope with defective algorithms and shape and calibrate their perceptions while working with AV. We found that, as front-line workers, safety drivers are forced to take risks accumulated from the AV industry upstream and are also confronting restricted self-development in working for AV development. We contribute the first empirical evidence of the lived experience of safety drivers, the first passengers in the development of AV, and also the grassroots workers for AV, which can shed light on future human-AI interaction research.

Recommended citation: M Chu, K Zong, X Shu, J Gong, Z Lu, K Guo, X Dai, G Zhou (2023). Work with AI and work for AI: autonomous vehicle safety drivers... lived experiences. Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems ....

Understanding Embodied Reference with Touch-Line Transformer. Permalink

Published in ICLR, 2023

Authors: Y Li, X Chen, H Zhao, Jiangtao Gong, G Zhou, F Rossano, Y Zhu

Abstract: We study embodied reference understanding, the task of locating referents using embodied gestural signals and language references. Human studies have revealed that objects referred to or pointed to do not lie on the elbow-wrist line, a common misconception; instead, they lie on the so-called virtual touch line. However, existing human pose representations fail to incorporate the virtual touch line. To tackle this problem, we devise the touch-line transformer: It takes as input tokenized visual and textual features and simultaneously predicts the referent's bounding box and a touch-line vector. Leveraging this touch-line prior, we further devise a geometric consistency loss that encourages the co-linearity between referents and touch lines. Using the touch-line as gestural information improves model performances significantly. Experiments on the YouRefIt dataset show our method achieves a +25.0% accuracy improvement under the 0.75 IoU criterion, closing 63.6% of the gap between model and human performances. Furthermore, we computationally verify prior human studies by showing that computational models more accurately locate referents when using the virtual touch line than when using the elbow-wrist line.

Recommended citation: Y Li, X Chen, H Zhao, J Gong, G Zhou, F Rossano, Y Zhu (2023). Understanding Embodied Reference with Touch-Line Transformer.. ICLR.

Human Understanding理解人22
Cognitive Instruments认知工具6

SVATA: A Spatial Visual Attention Tracking and Analysis Platform for Embodied Cognition Research Permalink

Published in Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems ..., 2026

Authors: X Ren, J Huang, S Feng, Y Wei, Jiangtao Gong, S Ma

Abstract: Understanding spatial visual attention is important for embodied cognition research, yet practical platforms for 3D attention analysis remain limited. We present SVATA—Spatial Visual Attention Tracking and Analysis, an open-source platform that supports an end-to-end workflow for collecting, analyzing, and visualizing world-referenced gaze-and-movement data within a 3D spatial context. SVATA maps multimodal signals (gaze, head, position) onto reconstructed geometry and computes a physiologically informed Average Focus Weight (\(AFW/m^{2}\)) metric as a proxy for overt visual focus. This representation supports structured analysis and multidimensional visualization of spatial viewing patterns. We evaluated SVATA through an in-the-wild museum deployment with 78 visitors and an expert study with seven prospective users; the results suggest its feasibility and perceived utility for analyzing spatial viewing behavior in embodied cognition research.

Recommended citation: X Ren, J Huang, S Feng, Y Wei, J Gong, S Ma (2026). SVATA: A Spatial Visual Attention Tracking and Analysis Platform for Embodied Cognition Research. Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems ....

An EEG dataset for understanding driving expertise from naturalistic urban road experiments Permalink

Published in Scientific Data, 2026

Authors: Jiangtao Gong, Y Yu, Y Cao, R Yang, X Chang, H Tang, X Zheng, Y Liu, S You, …

Abstract: Modern autonomous driving algorithms in complex urban environments strive for safety, comfort, and intelligence, qualities exemplified by expert drivers. This paper presents a comprehensive EEG dataset comparing 10 expert and 10 novice drivers in 13 naturalistic urban driving conditions on a fixed 5.7-kilometer route in in a city. Our multi-modal dataset synchronizes brain activity with vehicle CAN bus data, traffic participant information, and psychophysiological measures (Electrodermal Activity and Heart Rate) from drivers. Uniquely, we also collected physiological and subjective feedback from two passengers per trip to validate driving performance quality. All participants (drivers and passengers) completed a series of standardized subjective questionnaires at pre- and post-experiment. Finally, they participated in post-drive semi-structured interviews exploring driver decision-making processes and passenger experience. This novel dataset enables researchers to decode the neural signatures underlying driving expertise, providing valuable insights for developing more human-like, intelligent autonomous driving algorithms that can better navigate the complexities of urban traffic environments.

Recommended citation: J Gong, Y Yu, Y Cao, R Yang, X Chang, H Tang, X Zheng, Y Liu, S You,... (2026). An EEG dataset for understanding driving expertise from naturalistic urban road experiments. Scientific Data.

Human and algorithmic visual attention in driving tasks Permalink

Published in npj Artificial Intelligence 2 (1), 23, 2026

Authors: C Zheng, P Li, B Jin, S You, KI Chan, YQ Zhang, G Zhou, Jiangtao Gong

Abstract: Amid the heated debate on whether artificial intelligence possesses a human-like capacity for understanding, the compatibility and interaction between human and algorithmic visual attention remain unclear. Here, we address this issue through the lens of spatial and feature-based attention. Using autonomous driving as an epitome of safety-critical domains, we show that human attention in driving tasks can be divided into three phases, each characterized by spatial, feature-based, and mixed visual attention. Comparisons between each phase of human attention and algorithmic attention revealed a complex landscape of human-AI resemblance. For specialized detection and planning algorithms, incorporating semantic-rich, feature-based human attention markedly enhanced performance, suggesting these models lack human-like semantic visual understanding. In contrast, for large-scale Vision-Language Models, the effect was task-dependent, suggesting that while foundation models have bridged the “reasoning gap” through massive pre-training, a “grounding gap” persists in fine-grained visual tasks. Crucially, our findings demonstrate that incorporating human semantic attention offers an effective and economic pathway to compensating for these gaps, enhancing model understanding in safety-critical and grounding-heavy tasks without the need for massive scale.

Recommended citation: C Zheng, P Li, B Jin, S You, KI Chan, YQ Zhang, G Zhou, J Gong (2026). Human and algorithmic visual attention in driving tasks. npj Artificial Intelligence 2 (1), 23.

How Generative Music Affects the ISO Principle-Based Emotion-Focused Therapy: An EEG Study Permalink

Published in Proceedings of the Annual Meeting of the Cognitive Science Society 47, 2025

Authors: J Bao, Y Lyu, J Yang, Y Jin, Jiangtao Gong

Abstract: Recently, AI-generated content (AIGC) technologies have made remarkable advancements, even achieving superhuman performance across various domains. However, few previous studies have investigated its impact on emotion-focused therapy with artistic content, e.g., music. In this paper, we conducted an EEG experiment to explore the effects of generative music on emotion-focused music therapy based on the ISO principle. This experiment compared AI-generated and human-created music regarding the changes in participants' valence and arousal following negative emotion induction with the ISO principle adherence and non-adherence. The results show that generative music, with its harmonic consistency and simple rhythm, is more effective in supporting positive emotions and improving temporal lobe activity. Besides, the therapeutic effectiveness of generative music adhering to the ISO principle has also been validated. This study highlights the distinct emotional and neural mechanisms of AI-generated music, offering valuable insights into future AI-powered emotion-focused therapy strategies.

Recommended citation: J Bao, Y Lyu, J Yang, Y Jin, J Gong (2025). How Generative Music Affects the ISO Principle-Based Emotion-Focused Therapy: An EEG Study. Proceedings of the Annual Meeting of the Cognitive Science Society 47.

Distinct neural pathway and its information flow for blind individual’s Braille reading Permalink

Published in NeuroImage 300, 120852, 2024

Authors: R Wang, Jiangtao Gong, C Zhao, Y Xu, B Hong

Abstract: • Left-lateralized occipito-frontal functional connectivity correlated with individual Braille reading proficiency. • Increased bidirectional LOC-IFC information flow with predominant top-down communication emerged in a natural Braille reading task. • Greater top-down modulation contributed to higher Braille reading proficiency. • A ‘two-tale’ model of the LOC-IFC neural link better predicts individual differences in natural Braille reading among late blind people. Natural Braille reading presents significant challenges to the brain networks of late blind individuals, yet its underlying neural mechanisms remain largely unexplored. Using natural Braille texts in behavioral assessments and functional MRI, we sought to pinpoint the neural pathway and information flow crucial for Braille reading performance in late blind individuals. In the resting state, we discovered a unique neural connection between the higher-order ‘visual’ cortex, the lateral occipital cortex (LOC), and the inferior frontal cortex (IFC) in late blind individuals, but not in sighted controls. The left-lateralized LOC-IFC connectivity was correlated with individual Braille reading proficiency. Prolonged Braille reading practice led to increased strength of this connectivity. During a natural Braille reading task, bidirectional information flow between the LOC and the IFC was positively modulated, with a predominantly stronger top-down modulation from the IFC to the LOC. This stronger top-down modulation contributed to higher Braille reading proficiency. We thus proposed a two-predictor multiple regression model to predict individual Braille reading proficiency, incorporating both static connectivity and dynamic top-down communication between the LOC-IFC link. This work highlights the dual contributions of the occipito-frontal neural pathway and top-down cognitive strategy to superior natural Braille reading performance, offering guidance for training late blind individuals.

Recommended citation: R Wang, J Gong, C Zhao, Y Xu, B Hong (2024). Distinct neural pathway and its information flow for blind individual's Braille reading. NeuroImage 300, 120852.

Quantitative Studies定量研究7

“It Must Be Gesturing Towards Me”: Gesture-Based Interaction between Autonomous Vehicles and Pedestrians Permalink

Published in Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems ..., 2024

Authors: X Chang, Z Chen, X Dong, Y Cai, T Yan, H Cai, Z Zhou, G Zhou, Jiangtao Gong

Abstract: Interacting with pedestrians understandably and efficiently is one of the toughest challenges faced by autonomous vehicles (AVs) due to the limitations of current algorithms and external human-machine interfaces (eHMIs). In this paper, we design eHMIs based on gestures inspired by the most popular method of interaction between pedestrians and human drivers. Eight common gestures were selected to convey AVs’ yielding or non-yielding intentions at uncontrolled crosswalks from previous literature. Through a VR experiment (N1 = 31) and a following online survey (N2 = 394), we discovered significant differences in the usability of gesture-based eHMIs compared to current eHMIs. Good gesture-based eHMIs increase the efficiency of pedestrian-AV interaction while ensuring safety. Poor gestures, however, cause misinterpretation. The underlying reasons were explored: ambiguity regarding the recipient of the signal and whether the gestures are precise, polite, and familiar to pedestrians. Based on this empirical evidence, we discuss potential opportunities and provide valuable insights into developing comprehensible gesture-based eHMIs in the future to support better interaction between AVs and other road users.

Recommended citation: X Chang, Z Chen, X Dong, Y Cai, T Yan, H Cai, Z Zhou, G Zhou, J Gong (2024). " It Must Be Gesturing Towards Me": Gesture-Based Interaction between Autonomous Vehicles and Pedestrians. Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems ....

Can quadruped guide robots be used as guide dogs? Permalink

Published in 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems ..., 2023

Authors: L Wang, Q Chen, Y Zhang, Z Li, T Yan, F Wang, G Zhou, Jiangtao Gong

Abstract: Quadruped robots have the potential to guide blind and low vision (BLV) people due to their highly flexible locomotion and emotional value provided by their bionic forms. However, the development of quadruped guide robots rarely involves BLV users' participatory designs and evaluations. In this paper, we conducted two empirical experiments both in indoor controlled and outdoor field scenarios, exploring the benefits and drawbacks of quadruped guide robots. The results show that the nowadays commercial quadruped robots exposed significant disadvantages in usability and trust compared with wheeled robots. It is concluded that the moving gait and walking noise of quadruped robots would limit the guiding effectiveness to a certain extent, and the empathetic effect of its bionic form for BLV users could not be fully reflected. Based on the findings of wheeled robots and quadruped robots' advantages, we discuss the design implications for the future guide robot design for BLV users. This paper reports the first empirical experiment about quadruped guide robots with BLV users and preliminary explores their potential improvement space in substituting guide dogs, which can inspire the further specialized design of quadruped guide robots.

Recommended citation: L Wang, Q Chen, Y Zhang, Z Li, T Yan, F Wang, G Zhou, J Gong (2023). Can quadruped guide robots be used as guide dogs?. 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems ....

Side-by-side vs face-to-face: Evaluating colocated collaboration via a transparent wall-sized display Permalink

Published in Proceedings of the ACM on Human-Computer Interaction 7 (CSCW1), 1-29, 2023

Authors: Jiangtao Gong, J Sun, M Chu, X Wang, M Luo, Y Lu, L Zhang, Y Wu, Q Wang, …

Abstract: Traditional wall-sized displays mostly only support side-by-side co-located collaboration, while transparent displays naturally support face-to-face interaction. Many previous works assume transparent displays support collaboration. Yet it is unknown how exactly its afforded face-to-face interaction can support loose or close collaboration, especially compared to the side-by-side configuration offered by traditional large displays. In this paper, we used an established experimental task that operationalizes different collaboration coupling and layout locality, to compare pairs of participants collaborating side-by-side versus face-to-face in each collaborative situation. We compared quantitative measures and collected interview and observation data to further illustrate and explain our observed user behavior patterns. The results showed that the unique face-to-face collaboration brought by transparent display can result in more efficient task performance, different territorial behavior, and both positive and negative collaborative factors. Our findings provided empirical understanding about the collaborative experience supported by wall-sized transparent displays and shed light on its future design.

Recommended citation: J Gong, J Sun, M Chu, X Wang, M Luo, Y Lu, L Zhang, Y Wu, Q Wang,... (2023). Side-by-side vs face-to-face: Evaluating colocated collaboration via a transparent wall-sized display. Proceedings of the ACM on Human-Computer Interaction 7 (CSCW1), 1-29.

Can quadruped navigation robots be used as guide dogs? Permalink

Published in arXiv preprint arXiv:2210.08727, 2022

Authors: L Wang, Q Chen, Y Zhang, Z Li, T Yan, F Wang, G Zhou, Jiangtao Gong

Abstract: Quadruped robots have the potential to guide blind and low vision (BLV) people due to their highly flexible locomotion and emotional value provided by their bionic forms. However, the development of quadruped guide robots rarely involves BLV users' participatory designs and evaluations. In this paper, we conducted two empirical experiments both in indoor controlled and outdoor field scenarios, exploring the benefits and drawbacks of quadruped guide robots. The results show that the nowadays commercial quadruped robots exposed significant disadvantages in usability and trust compared with wheeled robots. It is concluded that the moving gait and walking noise of quadruped robots would limit the guiding effectiveness to a certain extent, and the empathetic effect of its bionic form for BLV users could not be fully reflected. Based on the findings of wheeled robots and quadruped robots' advantages, we discuss the design implications for the future guide robot design for BLV users. This paper reports the first empirical experiment about quadruped guide robots with BLV users and preliminary explores their potential improvement space in substituting guide dogs, which can inspire the further specialized design of quadruped guide robots.

Recommended citation: L Wang, Q Chen, Y Zhang, Z Li, T Yan, F Wang, G Zhou, J Gong (2022). Can quadruped navigation robots be used as guide dogs?. arXiv preprint arXiv:2210.08727.

“I can’t name it, but I can perceive it”: Conceptual and Operational Design of “Tactile Accuracy” Assisting Tactile Image Cognition Permalink

Published in Proceedings of the 22nd International ACM SIGACCESS Conference on Computers ..., 2020

Authors: Jiangtao Gong, W Yu, L Ni, Y Jiao, Y Liu, X Fu, Y Xu

Recommended citation: J Gong, W Yu, L Ni, Y Jiao, Y Liu, X Fu, Y Xu (2020). "I can't name it, but I can perceive it": Conceptual and Operational Design of "Tactile Accuracy" Assisting Tactile Image Cognition. Proceedings of the 22nd International ACM SIGACCESS Conference on Computers ....

影响触觉图像识别因素的定量分析 Permalink

Published in 计算机辅助设计与图形学学报 30 (8), 1438-1445, 2018

Authors: 龚江涛,于文凤,屈同希,刘烨,傅小兰,徐迎庆

Recommended citation: 龚江涛,于文凤,屈同希,刘烨,傅小兰,徐迎庆 (2018). 影响触觉图像识别因素的定量分析. 计算机辅助设计与图形学学报 30 (8), 1438-1445.

Multi-factor Analysis Assisting T-Image Design for Tactile Cognition Permalink

Published in Journal of Computer-Aided Design & Computer Graphics 30 (8), 1438-1445, 2018

Authors: Jiangtao Gong, W Yu, T Qu, Y Liu, X Fu, Y Xu

Abstract: 为了使更多盲人能受益于盲文书籍所伴随的插图,区别于传统的V图像(视觉图像),对设计适合触觉认知的T图像(触觉图像)提出新的设计原则.首先将242张常见物品的V图像制作为线条凸起的可触摸图片;然后邀请10位盲人被试和10位蒙眼明眼人被试通过触摸来尽量准确地命名这些线条图,并要求被试在触摸的过程中进行"出声思维";再根据被试对线条图的描述,提取22个可能影响二维线条图触觉识别的特征;最后以识别正确率作为图片识别难易程度的指标,使用随机森林算法进行了特征建模,并对所有特征进行单因素和多因素的回归分析.实验结果表明,通过随机森林算法建立的模型,可以基于图片中这些特征预测图片触觉识别的难易程度;通过多因素回归分析,提取出对触觉识别有显著影响力的几个重要特征,并用于指导T图像的设计.

Recommended citation: J Gong, W Yu, T Qu, Y Liu, X Fu, Y Xu (2018). Multi-factor Analysis Assisting T-Image Design for Tactile Cognition. Journal of Computer-Aided Design & Computer Graphics 30 (8), 1438-1445.

Qualitative Studies定性研究8

“I Will Dream Sweet Dreams”: Understanding Remote Companionship Volunteer Activities for Factual Orphans in Rural China Permalink

Published in Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems ..., 2026

Authors: J Zhang, Y Tang, Y Hu, T Wang, H Song, Z Lu, Jiangtao Gong

Abstract: A significant number of De facto orphans in underdeveloped regions face potential mental health risks, while some novel interventions attempt to involve university student volunteers in providing companionship. However, there is little documented literature on current practices of such emotional support volunteer activities. We conducted a comprehensive investigation of multiple stakeholders involved in an online companionship program facilitated by university student volunteers, designed to provide remote emotional support to these children. Through field observations and interviews, we summarize current practices and identify benefits. We discovered that current remote volunteer initiatives face numerous challenges impeding companionship effectiveness. We summarized four potential technological requirements derived from these challenges. Our study provides the first documented account of remote volunteer companionship activities aimed at improving vulnerable children’s mental health. Our research can inspire future developments of technological solutions for improving emotional companionship in volunteer programs and vulnerable children’s wellbeing.

Recommended citation: J Zhang, Y Tang, Y Hu, T Wang, H Song, Z Lu, J Gong (2026). "I Will Dream Sweet Dreams": Understanding Remote Companionship Volunteer Activities for Factual Orphans in Rural China. Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems ....

Exploring Wearable Design for Emotional Health and Well-Being during Menopause: Perspectives and Design Opportunities Permalink

Published in Proceedings of the Extended Abstracts of the CHI Conference on Human Factors ..., 2025

Authors: W Zhang, H Bao, C Li, Y Zhai, J Liu, Jiangtao Gong

Recommended citation: W Zhang, H Bao, C Li, Y Zhai, J Liu, J Gong (2025). Exploring Wearable Design for Emotional Health and Well-Being during Menopause: Perspectives and Design Opportunities. Proceedings of the Extended Abstracts of the CHI Conference on Human Factors ....

Understanding human-AI collaboration in music therapy through co-design with therapists Permalink

Published in Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems ..., 2024

Authors: J Sun, J Yang, G Zhou, Y Jin, Jiangtao Gong

Abstract: The rapid development of musical AI technologies has expanded the creative potential of various musical activities, ranging from music style transformation to music generation. However, little research has investigated how musical AIs can support music therapists, who urgently need new technology support. This study used a mixed method, including semi-structured interviews and a participatory design approach. By collaborating with music therapists, we explored design opportunities for musical AIs in music therapy. We presented the co-design outcomes involving the integration of musical AIs into a music therapy process, which was developed from a theoretical framework rooted in emotion-focused therapy. After that, we concluded the benefits and concerns surrounding music AIs from the perspective of music therapists. Based on our findings, we discussed the opportunities and design implications for applying musical AIs to music therapy. Our work offers valuable insights for developing human-AI collaborative music systems in therapy involving complex procedures and specific requirements.

Recommended citation: J Sun, J Yang, G Zhou, Y Jin, J Gong (2024). Understanding human-AI collaboration in music therapy through co-design with therapists. Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems ....

” I see it as a wellspring for my positive and upward journey in life.”: Understanding Current Practices of Assistive Technology’s Customized Modification in China Permalink

Published in Proceedings of the ACM on Human-Computer Interaction 8 (CSCW2), 1-42, 2024

Authors: K Yang, J Wu, H Xin, Jiangtao Gong

Abstract: Due to the significant differences in physical conditions and living environments of people with disabilities, standardized assistive technologies (ATs) often fail to meet their needs. Modified AT, especially DIY (Do It Yourself) ATs, are a popular solution in many high-income countries, but there is a lack of documentation for low- and middle-income areas, especially in China, where the culture of philanthropy is undeveloped. To understand the current situation in this paper, we conducted semi-structured interviews with 10 individuals with disabilities using modified ATs and 10 individuals involved in providing these including family members, standard assistive device manufacturers, and individuals employed for their modification skills, etc. Based on the results of the thematic analysis, we have summarized the general process of modified ATs for people with disabilities in China and the benefits these devices bring. We found that modified ATs not only make the lives of people with disabilities more comfortable and convenient but also bring them confidence, reduce social pressure, and even help them achieve self-realization. Additionally, we summarized the challenges they encountered before, during, and after the modification, including awareness gaps, family resistance, a lack of a business model, and so on. Specifically, we conducted a special case study about the typical business models and challenges currently faced by AT Modification Organizations in China. Our research provides important design foundations and research insights for the future of universal and personalized production of AT.

Recommended citation: K Yang, J Wu, H Xin, J Gong (2024). " I see it as a wellspring for my positive and upward journey in life.": Understanding Current Practices of Assistive Technology's Customized Modification in China. Proceedings of the ACM on Human-Computer Interaction 8 (CSCW2), 1-42.

Not a playroom, but a passage: Exploring the design space of a technology-mediated calm-down corner Permalink

Published in Extended Abstracts of the CHI Conference on Human Factors in Computing ..., 2024

Authors: J Gao, A Chen, Z Yi, X Yan, J Liu, G Zhou, Jiangtao Gong

Abstract: Preschool educators play an important role in supporting children’s emotion regulation capacities in ways that promote their socio-emotional and behavioral health. However, tracking children’s emotional changes and giving timely intervention to every child in the class simultaneously is challenging for preschool educators. In this paper, we use the "calm-down corner," which is widely used in classrooms with spaces where children can regain emotional control, as a probe to explore the design space of technology-supported preschool educators dealing with children’s emotional problems through one semi-structured interview with ten preschool educators and one co-design session with another 12 preschool educators. We summarized current challenges and the design considerations from preschool educators’ perspective, which could offer valuable insights for technology-enabled systems to facilitate the emotion regulation process for educators in the future.

Recommended citation: J Gao, A Chen, Z Yi, X Yan, J Liu, G Zhou, J Gong (2024). Not a playroom, but a passage: Exploring the design space of a technology-mediated calm-down corner. Extended Abstracts of the CHI Conference on Human Factors in Computing ....

Work with AI and work for AI: autonomous vehicle safety drivers… lived experiences Permalink

Published in Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems ..., 2023

Authors: M Chu, K Zong, X Shu, Jiangtao Gong, Z Lu, K Guo, X Dai, G Zhou

Abstract: The development of Autonomous Vehicle (AV) has created a novel job, the safety driver, recruited from experienced drivers to supervise and operate AV in numerous driving missions. Safety drivers usually work with non-perfect AV in high-risk real-world traffic environments for road testing tasks. However, this group of workers is under-explored in the HCI community. To fill this gap, we conducted semi-structured interviews with 26 safety drivers. Our results present how safety drivers cope with defective algorithms and shape and calibrate their perceptions while working with AV. We found that, as front-line workers, safety drivers are forced to take risks accumulated from the AV industry upstream and are also confronting restricted self-development in working for AV development. We contribute the first empirical evidence of the lived experience of safety drivers, the first passengers in the development of AV, and also the grassroots workers for AV, which can shed light on future human-AI interaction research.

Recommended citation: M Chu, K Zong, X Shu, J Gong, Z Lu, K Guo, X Dai, G Zhou (2023). Work with AI and work for AI: autonomous vehicle safety drivers... lived experiences. Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems ....

Remote co-teaching in rural classroom: Current practices, impacts, and challenges Permalink

Published in Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems ..., 2022

Authors: S Guo, T Sun, Jiangtao Gong, Z Lu, L Zhang, Q Wang

Abstract: The shortage of high-quality teachers is one of the biggest educational problems faced by underdeveloped areas. With the development of information and communication technologies (ICTs), China has begun a remote co-teaching intervention program using ICTs for rural classes, forming a unique “co-teaching classroom”. We conducted semi-structured interviews with nine remote urban teachers and twelve local rural teachers. We identified the remote co-teaching classes’ standard practices and co-teachers’ collaborative work process. We also found that remote teachers’ high-quality class directly impacted local teachers and students. Furthermore, interestingly, local teachers were also actively involved in making indirect impacts on their students by deeply coordinating with remote teachers and adapting the resources offered by the remote teachers. We conclude by summarizing and discussing the challenges faced by teachers, lessons learned from the current program, and related design implications to achieve a more adaptive and sustainable ICT4D program design.

Recommended citation: S Guo, T Sun, J Gong, Z Lu, L Zhang, Q Wang (2022). Remote co-teaching in rural classroom: Current practices, impacts, and challenges. Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems ....

All in one group: Current practices, lessons and challenges of Chinese home-school communication in IM Group Chat Permalink

Published in Proceedings of the 2021 CHI conference on human factors in computing systems ..., 2021

Authors: Jiangtao Gong, Z Yao, Z Lu, Q Ding, Y Zhang, L Zhang, Q Wang

Abstract: When schools and families form a good partnership, children benefit. With the recent flourishing of communication apps, families and schools in China have shifted their primary communication channels to chat groups hosted on popular instant-messenger(IM) tools such as WeChat and QQ. With an interview study consisting of 18 parents and 9 teachers, followed by a survey study with 210 teachers, we found that IM group chat has become the most popular way that the majority of parents and teachers communicate, from among the many different channels available. While there are definite advantages to this kind of group chat, we also found a number of problematic issues, including a lack of privacy and repeated negative feedback shared by both parents and teachers. We discuss our results on how IM-based group chat could affect Chinese teachers’ authoritative figures, affect Chinese teacher’s work-life balance and potentially compromise Chinese students’ privacy.

Recommended citation: J Gong, Z Yao, Z Lu, Q Ding, Y Zhang, L Zhang, Q Wang (2021). All in one group: Current practices, lessons and challenges of Chinese home-school communication in IM Group Chat. Proceedings of the 2021 CHI conference on human factors in computing systems ....

Surveys综述6

Evaluating Text-based Conversational Agents for Mental Health: A Systematic Review of Metrics, Methods and Usage Contexts Permalink

Published in Proceedings of the 2025 International Conference on Human-Engaged Computing ..., 2025

Authors: Jiangtao Gong, X Wen, F Tao, X Wang, X Yang, Y Tang

Abstract: Text-based conversational agents (CAs) are increasingly used in mental health, yet evaluation practices remain fragmented. We conducted a PRISMA-guided systematic review (May–June 2024) across ACM Digital Library, Scopus, and PsycINFO. From 613 records, 132 studies were included, with dual-coder extraction achieving substantial agreement (κ = 0.77–0.92). We synthesized evaluation approaches across three dimensions: metrics, methods, and usage contexts. Metrics were classified into CA-centric attributes (e.g., reliability, safety, empathy) and user-centric outcomes (experience, knowledge, psychological state, health behavior). Methods included automated analyses, standardized psychometric scales, and qualitative inquiry. Temporal designs ranged from momentary to follow-up assessments. Findings show reliance on Western-developed scales, limited cultural adaptation, predominance of small and short-term samples, and weak links between automated performance metrics and user well-being. We argue for methodological triangulation, temporal rigor, and equity in measurement. This review offers a structured foundation for reliable, safe, and user-centered evaluation of mental health CAs.

Recommended citation: J Gong, X Wen, F Tao, X Wang, X Yang, Y Tang (2025). Evaluating Text-based Conversational Agents for Mental Health: A Systematic Review of Metrics, Methods and Usage Contexts. Proceedings of the 2025 International Conference on Human-Engaged Computing ....

Human-computer interaction for virtual-real fusion Permalink

Published in Journal of Image and Graphics 28 (6), 1513-1542, 2023

Authors: J Tao, Jiangtao Gong, N Gao, S Fu, S Liang, C Yu

Abstract: 面向虚实融合的人机交互涉及计算机科学、认知心理学、人机工程学、多媒体技术和虚拟现实等领域,旨在提高人机交互的效率,同时响应人类认知与情感的需求,在办公教育、机器人和虚拟/增强现实设备中都有广泛应用。本文从人机交互涉及感知计算、人与机器人交互及协同、个性化人机对话和数据可视化等4个维度系统阐述面向虚实融合人机交互的发展现状。对国内外研究现状进行对比,展望未来的发展趋势。本文认为兼具可迁移与个性化的感知计算、具备用户行为深度理解的人机协同、用户自适应的对话系统等是本领域的重要研究方向。;Virtual-real human-computer interaction(VR-HCI) is an interdisciplinary field that encompasses human and computer interactions to address human-related cognitive and emotional needs. This interdisciplinary knowledge integrates domains such as computer science, cognitive psychology, ergonomics, multimedia technology, and virtual reality. With the advancement of big data and artificial intelligence, VR-HCI benefits industries like education, healthcare, robotics, and entertainment, and is increasingly recognized as a key supporting technology for metaverse-related development. In recent years, machine learning-based human cognitive and emotional analysis has evolved, particularly in applications like robotics and wearable interaction devices. As a result, VR-HCI has focused on the challenging issue of creating intelligent and anthropomorphic interaction systems. This literature review examines the growth of VR-HCI from four aspects:perceptual computing, human-machine interaction and coordination, human-computer dialogue interaction, and data visualization. Perceptual computing aims to model human daily life behavior, cognitive processes, and emotional contexts for personalized and efficient human-computer interactions. This discussion covers three perceptual aspects related to pathways, objects, and scenes. Human-machine interaction scenarios involve virtual and real-world integration and perceptual pathways, which are divided into primary perception types:visual-based, sensor-based, and wireless non-contact. Object-based perception is subdivided into personal and group contexts, while scene-based perception is subdivided into physical behavior and cognitive contexts. Human-machine interaction primarily encompasses technical disciplines such as mechanical and electrical engineering, computer and control science, artificial intelligence, and other related arts or humanistic disciplines like psychology and design. Human-robot interaction can be categorized by functional mechanisms into 1) collaborative operation robots, 2) service and assistance robots, and 3) social, entertainment, and educational robots. Key modules in human-computer dialogue interaction systems include speech recognition, speaker recognition, dialogue system, and speech synthesis. The level of intelligence in these interaction systems can be further enhanced by considering users' inherent characteristics, such as speech pronunciation, preferences, emotions, and other attributes. For human-machine interaction, it mainly involves technical disciplines in relevant to mechanical and electrical engineering, computer and control science, and artificial intelligence, as well as other related arts or humanistic disciplines like psychology and design. Humans-robots interaction can be segmented into three categories in terms of its functional mechanism:1) collaborative operation robots, 2) service and assistance robots, and 3) social, entertainment and educational robots. For human-computer dialogue interaction, the system consists of such key modules like speech recognition, speaker recognition, dialogue system, and speech synthesis. The microphone sensor can pick up the speech signal, which is then converted to text information through the speech recognition module. The dialogue system can process the text information, understand the user's intention, and generates a reply. Finally, the speech-synthesized module can convert the reply information into speech information, completing the interaction process. In recent years, the level of intelligence of the interaction system can be further improved by combining users' inherent characteristics such as speech pronunciation, preferences, emotions, and other characteristics, optimizing the various modules of the interaction system. For data transformation and visualization, it is benched for performing data cleaning tasks on tabular data, and various tools in R and Python can perform these tasks as well. Many software systems have developed graphical user interfaces to assist users in completing data transformation tasks, such as Microsoft Excel, Tableau Prep Builder, and OpenRefine. Current recommendation-based algorithms interactive systems are beneficial for users transform data easily. Researchers have also developed tools that can transform network structures. We analyze its four aspects of 1) interactive data transformation, 2) data transformation visualization, 3) data table visual comparison, and 4) code visualization in human-computer interaction systems. We identify several future research directions in VR-HCI, namely 1) designing generalized and personalized perceptual computing, 2) building human-machine cooperation with a deep understanding of user behavior, and 3) expanding user-adaptive dialogue systems. For perceptual computing, it still lacks joint perception of multiple devices and individual differences in human behavior perception. Furthermore, most perceptual research can use generalized models, neglecting individual differences, resulting in lower perceptual accuracy, making it difficult to apply in actual settings. Therefore, future perceptual computing research trends are required for multimodal, transferable, personalized, and scalable research. For human-machine interaction and coordination, a systematic approach is necessary for constructing a design for human-machine interaction and collaboration. This approach requires in-depth research on user understanding, construction of interaction datasets, and long-term user experience. For human-computer dialogue interaction, current research mostly focuses on open-domain systems, which use pre-trained models to improve modeling accuracy for emotions, intentions, and knowledge. Future research should be aimed at developing more intelligent human-machine conversations that cater to individual user needs. For data transformation and visualization in HCI, the future directions can be composed of two parts:1) the intelligence level of data transformation can be improved through interaction for individual data workers on several aspects, e. g., appropriate algorithms for multiple types of data, recommendations for consistent user behavior and real-time analysis to support massive data. 2) The focus is on the integration of data transformation and visualization among multiple users, including designing collaborative mechanisms, resolving conflicts in data operation, visualizing complex data transformation codes, evaluating the effectiveness of various visualization methods, and recording and displaying multiple human behaviors. In summary, the development of VR-HCI can provide new opportunities and challenges for human-computer interaction towards Metaverse, which has the potential to seamlessly integrate virtual and real worlds.

Recommended citation: J Tao, J Gong, N Gao, S Fu, S Liang, C Yu (2023). Human-computer interaction for virtual-real fusion. Journal of Image and Graphics 28 (6), 1513-1542.

Human-computer interaction for virtual-reality fusion Permalink

Published in Chinese Journal of Image and Graphics 28 (6), 1513-1542, 2023

Authors: T Jianhua, G Jiangtao, G Nan, F Siwei, L Shan, Y Chun, Jiangtao Gong

Recommended citation: T Jianhua, G Jiangtao, G Nan, F Siwei, L Shan, Y Chun (2023). Human-computer interaction for virtual-reality fusion. Chinese Journal of Image and Graphics 28 (6), 1513-1542.

Classification, application, challenge, and future of midair gestures in augmented reality Permalink

Published in Journal of Sensors 2022 (1), 3208047, 2022

Authors: Y Lu, X Wang, Jiangtao Gong, L Zhou, S Ge

Abstract: Augmented Reality (AR) technology provides many opportunities to enhance people’s experience in interacting with data. Midair gesture, a natural interaction mode in AR, interacts with virtual elements without any auxiliary devices. It has become a hot topic of interest for researchers with the development of gesture recognition technology. From the perspective of user experience, the types of midair gesture, gesture recognition technology, applications were discussed. The challenges of air gesture interaction from two aspects of user experience and technology were analyzed, and the importance of collaborative gesture and low-cost user experience in the future were emphasized. Finally, the application prospect of air gesture interaction distance education, medical health, industry, office, and so on is discussed.

Recommended citation: Y Lu, X Wang, J Gong, L Zhou, S Ge (2022). Classification, application, challenge, and future of midair gestures in augmented reality. Journal of Sensors 2022 (1), 3208047.

触觉二维图像识别的认知机制 Permalink

Published in 心理科学进展 27 (4), 611, 2019

Authors: 于文凤,刘烨,傅小兰,龚江涛,徐迎庆

Abstract: The two-dimension tactile image is the main approach of translating visual information into haptic information. It plays an important role in helping visually impaired people perceive the external world. The recognition of haptic two-dimension image is considered to be based on the “visual translation” process where the haptic input is translated into the visual image. This process is influenced by the graphic geometric feature, perspective, visual experience, capability of visual representation, the process of tactile exploration, training and age. Exploration of the cognitive neural mechanism of two-dimension images haptic recognition is significant for improving the design and usability of two-dimension tactile images.

Recommended citation: 于文凤,刘烨,傅小兰,龚江涛,徐迎庆 (2019). 触觉二维图像识别的认知机制. 心理科学进展 27 (4), 611.