Instructions to use py-feat/face_multitask_v2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Py-Feat
How to use py-feat/face_multitask_v2 with Py-Feat:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
face_multitask_v2
Licensing scope: See the license and provenance notice before relying on this card's license metadata for pretrained-weight redistribution or commercial use. Existing valid grants are preserved.
A single multi-task convolutional model for facial behavior analysis, used by
py-feat's Detectorv2. From one face crop
it jointly predicts action units, categorical emotion, valence/arousal,
eye gaze, a 478-point face mesh, 6-DoF head pose, and 52 MediaPipe/ARKit
blendshapes.
Current default: face_multitask_v28.safetensors.
- Backbone: ConvNeXt-V2 Tiny (FCMAE + IN-22k/IN-1k pretrained)
- Heads: ME-GraphAU AU graph (AFG/FGG/SC) + unified-feature emotion/V-A and gaze heads + landmark, pose, and blendshape regression heads
- Params: ~30M · Input: 224×224 RGB (from a 256×256 face crop)
- Format: safetensors, with the
ModelV2ConfigJSON in the file metadata - Weights: uniform weight average ("model soup") of three training epochs
Older files in this repo (face_multitask_v2, _v26, _v27) are retained so
existing installs keep working. Each py-feat release pins the filename it was
built against — older code cannot construct newer architectures.
Outputs
| Task | Output | Notes |
|---|---|---|
| Action Units | 20 probabilities [0,1] | AU01,02,04,05,06,07,09,10,11,12,14,15,17,20,23,24,25,26,28,43 |
| Emotion | 7-class softmax | Neutral, Happy, Sad, Surprise, Fear, Disgust, Anger |
| Valence / Arousal | 2 × [−1,1] | tanh |
| Gaze | (yaw, pitch) radians | RAW convention is y-down: yaw+ = subject's right (image-left), pitch+ = looking DOWN. Detectorv2 negates pitch so Fex columns are canonical +up (since py-feat 2.1.1) |
| Face mesh | 478 × (x,y,z) | MediaPipe topology, chip-pixel coords (z = relative depth) |
| Head pose | (pitch, yaw, roll, tx, ty, tz) | radians / pixels; RAW pitch+ = down (img2pose teacher frame); Detectorv2 outputs canonical +up (since py-feat 2.1.1) |
| 68 landmarks | derived | dlib-68 subset sampled from the 478 mesh |
| Blendshapes | 52 coefficients [0,1] | MediaPipe/ARKit standard names |
Evaluation protocol
Every benchmark below is held out from training at the split level, and
DISFA+ additionally at the identity level. The training run reserves 16
splits: affectnet:val, raf_db:test, ferplus:{test,val},
meld:{bench,val}, afew_va:val, aff_wild2:bench, aff_wild2_va:val,
gaze360:{bench,val}, mpii_gaze:test, mpii_facegaze:test,
columbia_gaze:test, ethxgaze:test, eyediap:test. The disfaplus and
mpii_facegaze sources are excluded outright, and DISFA+ identities appearing
in other corpora are excluded as well.
Benchmarks
Chip-protocol inference on held-out splits.
| Task | Dataset | Metric | Score |
|---|---|---|---|
| AU | DISFA+ (12-AU, Cheong protocol) | macro-F1 | 0.682 |
| AU | DISFA+ (common-8 subset) | macro-F1 | 0.772 |
| Emotion | AffectNet val (7-cls, drop Contempt) | acc / macro-F1 | 0.628 / 0.625 |
| Emotion | RAF-DB official test (7-cls) | acc / macro-F1 | 0.869 / 0.806 |
| Valence/Arousal | AffectNet val | CCC (V / A) | 0.773 / 0.650 |
| Valence/Arousal | Aff-Wild2 official validation | CCC (V / A) | 0.376 / 0.477 |
| Valence/Arousal | AFEW-VA validation | CCC (V / A) | 0.687 / 0.548 |
| Gaze | Gaze360 (held-out split) | mean angular err | 13.04° |
| Gaze | MPIIGaze (leave-subject-out) | mean angular err | 8.26° |
| Gaze | ETH-XGaze (test) | mean angular err | 4.76° |
| Gaze | Columbia (test) | mean angular err | 3.76° |
| Gaze | EYEDIAP (test, never trained on) | mean angular err | 10.61° |
Cross-tool, end-to-end
Full shipped pipeline (detect → align → predict) on raw frames, scored on the images all tools processed. AU presence = DISFA+ intensity ≥ 2, prediction ≥ 0.5.
| Benchmark | This model | OpenFace 3.0 | LibreFace |
|---|---|---|---|
| DISFA+ AU, common-8 macro-F1 | 0.774 | 0.732 | 0.492 |
| DISFA+ AU, 12-AU macro-F1 | 0.671 | — (8 AUs only) | 0.397 |
| AffectNet-7 accuracy | 0.632 | 0.587 | 0.458 |
| AffectNet-7 macro-F1 | 0.632 | 0.587 | 0.410 |
| RAF-DB-7 accuracy | 0.880 | 0.673 | 0.746 |
| RAF-DB-7 macro-F1 | 0.815 | 0.586 | 0.580 |
AU robustness — perturbed DISFA+ (Cheong 2023 protocol)
Same 57,150 aligned DISFA+ crops as the main AU benchmark, with black-bar occlusion and luminance shifts. The 8-AU column uses the cross-tool common set (AU01/02/04/06/09/12/25/26).
| Perturbation | 12-AU macro-F1 | common-8 macro-F1 |
|---|---|---|
| none (baseline) | 0.685 | 0.774 |
| eyes occluded | 0.542 | 0.667 |
| mouth occluded | 0.538 | 0.632 |
| nose occluded | 0.678 | 0.785 |
| brightened | 0.593 | 0.684 |
| darkened | 0.675 | 0.765 |
Mouth occlusion is the worst case (−0.147 on 12-AU vs baseline).
Known limitations
- AU20 is effectively non-functional (per-AU F1 0.057 against a 0.682 macro). AU15 (0.479) and AU06 (0.526) are also well below the macro average. Do not rely on lip-stretch (AU20) predictions; treat AU15/AU06 with caution.
- Gaze pitch is uneven across domains. Per-axis pitch MAE is 3.06° on ETH-XGaze and 7.05° on EYEDIAP, but 5.21° on MPIIGaze with a lower prediction/ground-truth correlation (r = 0.755) — frontal, screen-directed gaze has the narrowest pitch range and is the weakest case. Yaw is consistently stronger than pitch on all three. Validate before depending on absolute gaze pitch in a frontal-camera setting.
- Landmarks are a secondary output. The 68 points are sampled from the 478 mesh rather than predicted by a dedicated landmark head, and 300W NME is correspondingly weaker than tools with a native 68-point head.
- Trained on posed and in-the-wild adult face imagery; performance on children, heavy occlusion, or extreme pose is not characterized.
Usage
from feat import Detectorv2
detector = Detectorv2(device="cuda")
fex = detector.detect("image.jpg") # returns a py-feat Fex
To pin this specific checkpoint:
detector = Detectorv2(device="cuda",
multitask_weights="face_multitask_v28.safetensors")
The model expects a face crop produced by RetinaFace + py-feat's
extract_face_from_bbox_torch(frame, bbox, face_size=256, expand_bbox=1.2),
then center-cropped to 224 and ImageNet-normalized. Detectorv2 handles this.
License and provenance
The weights have been designated for noncommercial research. Py-Feat's
original implementation is MIT, but the initialized
convnextv2_tiny.fcmae_ft_in22k_in1k weights
are CC BY-NC 4.0, distinct from the backbone's MIT source code.
For face_multitask_v28.safetensors, DISFA+ and EYEDIAP are excluded from
gradient training; their benchmark scores were used in checkpoint selection.
DISFA (without +) is training data. Training also includes BP4D, BP4D+, CK+,
UNBC-McMaster PAIN, EmotioNet, AM-FED, Aff-Wild2, AffectNet, RAF-DB, FER+,
ExpW, MELD, AFEW-VA, ETH-XGaze, Gaze360, MPIIGaze, Columbia Gaze, and
CelebV-HQ across the training stages. Older checkpoints have different
provenance and must be assessed separately.
For the current v2.8 checkpoint, Gaze360's research license, section 2(3) expressly excludes commercial applications of models trained on its dataset. ETH-XGaze access conditions expressly prohibit training models for commercial products. These are express restrictions, additional to the pretrained backbone's terms; they are not assumptions that all dataset licenses automatically transfer to weights. Gaze360's use and redistribution limits and ETH-XGaze's additional dataset/software terms also require separate review of public weight-sharing permission. The applicable acquisition terms and checkpoint provenance must be established separately for older files in this repository.
See the v2 licensing notice and dataset register. The research designation does not establish unrestricted public redistribution: the relevant dataset agreements, chronology, and any required permissions must be resolved for each checkpoint. This notice does not grant third-party rights or assume that all dataset or teacher terms automatically attach to every trained artifact. The optional ArcFace identity branch has its own terms and is not the basis for this network's licensing restrictions.
This notice clarifies scope; it does not revoke existing valid grants or create a new license for third-party material. Dataset and teacher terms do not automatically relicense every trained artifact or inference output. A software license or model-card badge alone does not establish all checkpoint redistribution or commercial-use permissions.
Commercial-use inquiries
These multitask weights are not offered as commercially cleared weights. Py-Feat does not broker third-party permissions or offer a commercial sublicense overriding upstream terms. Users can use independently cleared weights or contact the applicable rights holders directly for any necessary permissions covering the exact assets and commercial use. Dataset download approval alone is not necessarily such permission. See the v2 commercial-use notice and commercial retraining table. Existing valid grants remain unchanged; downstream permissions do not automatically resolve the distributor's separate obligations.