Training Output

#32 · closed · 3 comments

View on GitHub ↗

emlcpfx

![image](https://github.com/user-attachments/assets/395603ae-f437-4b8c-887c-041f4c31e963) What is the file format that gets outputted from this? Do you have an example image and data pair you could share?

Comments

georgeretsi

Hi there, the output is simply the trained weights for the 3 encoders. These are saved in the log_path. The training data are the same datasets for the smirk pipeline, but only the predicted landmarks and the predicted MICA shape parameters are used as targets. The existing code fully support this.

emlcpfx

Does that mean it’s virtually impossible to train this to handle extreme profile or even head turned almost backwards? Because if MICA can’t place the head in those situations, then you can’t generate the training data?

georgeretsi

Extreme profiles are not supported in this pipeline simply because the datasets used do not have extreme profiles. The focus of this work was expressions - and for expression typical poses make more sense. Of course one can use different datasets to include also extreme poses, as well as sota pose estimators to distill their performance into the SMIRK pipeline. If for some reason MICA parameters cannot be estimated (in extreme poses as you mentioned), you can always zero-out the corresponding loss. To do this a confidence of prediction is needed. These extra tools are not provided in the current framework but can be borrowed from relevant approaches, if needed.