Abstract
Multimodal interfaces increasingly incorporate social cues, such as self-disclosure, to enhance user engagement. However, how modality and disclosure depth shape users’ emotional and cognitive responses remains unclear. This exploratory study examines self-disclosure presented through text and voice across four conditions varying in modality and disclosure level. Physiological responses were measured as indicators of emotional engagement and relative changes in cognitive load. Text-only variations and adding voice to low self-disclosure produced no significant differences, whereas combining voice with high self-disclosure was associated with significantly heightened emotional engagement and a relative reduction in cognitive load. These exploratory results provide an initial empirical basis for future research on how voice and self-disclosure may work together in shaping user responses to socially expressive interfaces.

