To bridge this gap, researchers developed MindTS-MMD, a Chinese multimodal dataset designed to capture emotional expressions and tic-related behaviors in this specific pediatric demographic.
Understanding the MindTS-MMD Dataset Structure
The dataset contains 10,034 instance-level samples sourced from 60 children with Tourette syndrome. Participants ranged in age from six to 12 years old. Data collection relied on semi-structured emotion-elicitation tasks designed to prompt natural behavioral and emotional responses.
Each recorded sample integrates available symbolic representations across three distinct modalities: visual, acoustic, and semantic. Every instance links directly to a corresponding annotation record. This multi-layered approach ensures researchers have access to synchronized data streams for deeper computational analysis.
Categories and Annotation Metrics
Annotations within the dataset cover seven distinct emotional and behavioral categories: anxious, calm, focused, irritable, relaxed, shy, and tense. Beyond these categorical labels, the documentation includes child-reported and experimenter-observed valence-arousal ratings.
Crucially for clinical researchers, the annotations also track tic occurrence, anatomical location, and frequency whenever those movements are observable during the tasks. Developers evaluated several core metrics during compilation, including annotation reliability, signal quality, audiovisual synchronization, facial tracking precision, modality completeness, and tic-emotion co-occurrence.
Computational Usability and Baseline Models
To demonstrate the practical application of the released representations, the dataset providers included both unimodal and multimodal baselines. These baselines give engineers and data scientists a benchmark for evaluating how well machine learning models process the complex interplay between tics and emotional expressions.
Frequently Asked Questions
What is the primary purpose of the MindTS-MMD dataset?
The dataset supports computational research on emotional expression and tic-related behavior specifically in children with Tourette syndrome.
How many participants are included in the dataset?
The resource includes data collected from 60 children with Tourette syndrome, aged six to 12 years old, resulting in 10,034 instance-level samples.
What modalities are represented in the samples?
Each sample includes symbolic representations from three modalities: visual, acoustic, and semantic, alongside detailed annotation records.
What specific tic-related data is annotated?
Annotations track tic occurrence, anatomical location, and frequency whenever these behaviors are observable during the emotion-elicitation tasks.
Join the Discussion: How do you see multimodal datasets shaping the future of pediatric clinical research? Share your thoughts in the comments below, explore our related articles on computational health, and subscribe to our newsletter for the latest updates.
Worth a look