Why Subtle Emotion Recognition Must Be Personalized
Digital behaviour-change interventions increasingly use virtual coaches and online platforms to provide accessible and continuous support. Imagine a virtual health coach asking someone if they are ready to exercise more or follow a treatment plan. Although the person says “yes,” they pause, look away, and tighten their lips. A healthcare professional might recognize these subtle signs of uncertainty. A conventional digital system, however, might focus solely on the spoken response or recognize only broad emotions such as joy, sadness, and anger. Fine-grained states such as pain, stress, ambivalence, and hesitancy are harder to identify because their signals are often weaker, shorter, and expressed gradually over time.
Recognizing these subtle emotions is further amplified by individual differences. A raised eyebrow or tight lip may indicate discomfort in one person but be part of a neutral expression in another. Consequently, a model trained on many individuals may learn generic emotional patterns that do not accurately represent a new unseen person.
To address this challenge, CLIP-AUTT—a test-time personalization method for fine-grained emotion recognition— was developed at ÉTS. The method identifies the most informative temporal segment of an unlabeled video and adjusts a small set of facial-action descriptions to reflect the expressive patterns of an unseen person. This lightweight approach enables subject-specific adaptation without requiring emotion labels or a complete retraining of the model.
From Fixed Prompts to Personalized Emotion Recognition
Vision–language models (VLMs), such as CLIP, connect visual content with textual descriptions. For emotion recognition, they compare facial images or videos with descriptions such as an “expression of [Pain]” or “of [Stress].” As illustrated in Figure 1(a), conventional CLIP-based video emotion recognition methods rely on predefined emotion descriptions or prompts generated by large language models. The model is then trained or fine-tuned to connect videos with these descriptions. However, the prompts remain fixed when the model encounters a new person. Consequently, the same emotion descriptions are used for everyone, even though subtle expressions can differ considerably between individuals. Generic or automatically generated prompts may also be too broad to represent the small facial movements associated with pain, stress, ambivalence, or hesitancy.
To address these limitations, CLIP-AU was proposed as shown in Figure 1(b). Instead of relying on generic emotion descriptions, it uses prompts describing localized facial movements, known as action units (AUs) It compares these descriptions with facial movements across video frames and learns how combinations of AUs evolve in time. This provides a more detailed and interpretable representation of subtle expressions without requiring AU labels, language-model-generated prompts, or fine-tuning of the complete CLIP model.
CLIP-AUTT, shown in Figure 1(c), extends this approach through test-time personalization. When a video from an unseen person is received, the method identifies the temporal window with the clearest expressive cues and adapts only the AU prompts to that individual. The global visual model remains unchanged, keeping the adaptation lightweight. This approach not only improves recognition performance but also remains computationally efficient, maintaining a high processing speed at a low computational cost. As illustrated in Figure 1(d), CLIP-AUTT adjusts its understanding to how the current individual expresses the emotion.
How CLIP-AUTT Personalizes an Unseen Individual
Subtle facial expressions do not remain equally visible throughout a video. They often emerge gradually, reach a brief peak, and then fade. CLIP-AUTT divides each video into short, overlapping temporal windows, rather than treating all frames as equally informative. A temporal module analyzes the facial movements within each window and compares them with the AU descriptions learned by CLIP-AU.
The method then estimates how confidently each window corresponds to a consistent combination of AUs. A window with high uncertainty may contain a neutral face, a transition between expressions, or weak facial movements. CLIP-AUTT selects the window with the lowest uncertainty, which is considered the most informative segment for personalization. This prevents the model from adapting to ambiguous or uninformative parts of the video.
Using the selected window, CLIP-AUTT adjusts only the textual representations of the AUs. The visual encoder, temporal module, and emotion classifier remain unchanged. This makes personalization computationally efficient and reduces the risk of altering the general knowledge learned during training. The adaptation requires neither emotion labels for the new video nor access to the original training data. Once the prediction is completed, the adapted prompts are reset before the next video is processed.
Personalization Improves Fine-Grained Emotion Recognition
CLIP-AUTT was evaluated on three video datasets representing different fine-grained emotions: BioVid for pain, StressID for stress, and BAH for ambivalence and hesitancy. The evaluation considered videos from individuals who were not seen during training and compared CLIP-AUTT with existing CLIP-based and test-time adaptation methods.
CLIP-AUTT achieved a weighted accuracy of 81.5% on BioVid, 80.8% on StressID, and 69.8% on BAH. The strongest comparison method obtained 76.1%, 75.9%, and 67.9%, respectively. CLIP-AUTT therefore improved recognition by 5.4 percentage points for pain, 4.9 points for stress, and 1.9 points for ambivalence and hesitancy.
The improvement was also visible in the F1 score, which reflects how reliably the model distinguishes the different emotion levels or classes. On StressID, for example, CLIP-AUTT increased the F1 score from 59.4% to 77.9%. These results show that selecting an informative temporal segment and adapting AU prompts can provide more reliable predictions than applying the same fixed representation to every individual.
Conclusion
The CLIP-AUTT method developed at ÉTS offers a lightweight, efficient, and practical approach to personalizing fine-grained emotion recognition in videos. By combining AUs prompting with informative temporal-window selection, it adapts to individual expressive patterns without labels or complete model retraining, bringing digital health interventions closer to understanding and responding to each person’s subtle emotional cues.
Additional Information
Further details about the research presented in this article are available in the following paper: Zeeshan, Muhammad Osama, et al. “CLIP-AUTT: Test-Time Personalization with Action Unit Prompting for Fine-Grained Video Emotion Recognition.” European Conference on Computer Vision (ECCV), 2026.
References
Zeeshan, M. O., Sharafi, M., Savary, B., Koerich, A. L., Pedersoli, M., & Granger, E. (2026). Clip-autt: Test-time personalization with action unit prompting for fine-grained video emotion recognition. European Conference on Computer Vision (ECCV 2026)
González-González, M., Belharbi, S., Zeeshan, M., Sharafi, M., Aslam, M. H., Pedersoli, M., ... & Granger, E. (2026, April). BAH dataset for ambivalence/hesitancy recognition in videos for digital behavioural change. In International Conference on Learning Representations (Vol. 2026, pp. 41540-41586).
González-González, M., Belharbi, S., Zeeshan, M. O., Sharafi, M., Aslam, M. H., Sia, L., ... & Granger, E. (2026). Multimodal Ambivalence/Hesitancy Recognition in Videos for Personalized Digital Health Interventions. Affective Computing and Intelligent INteraction 2026.
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., ... & Sutskever, I. (2021, July). Learning transferable visual models from natural language supervision. In International conference on machine learning (pp. 8748-8763). PmLR.
Martinez, B., Valstar, M. F., Jiang, B., & Pantic, M. (2017). Automatic analysis of facial actions: A survey. IEEE transactions on affective computing, 10(3), 325-347.