| license | mit | ||||
|---|---|---|---|---|---|
| tags |
|
Model: huggingface.co/Isa0/cat-detection Dataset: CAT Dataset – Kaggle
Detects 9 facial landmarks on cats (eyes, ears, nose, mouth) and draws a bounding box around them.
Two variants, both mobile-friendly (well under 100 ms per inference on a phone):
| Variant | Backbone | Input size | ONNX file |
|---|---|---|---|
small |
MobileNetV3-Small | 224×224 | cat_landmark_small.onnx |
large |
MobileNetV3-Large | 256×256 | cat_landmark_large.onnx |
Both models take RGB input in [0, 1] (ImageNet normalization is baked into the
exported graph) and output 18 values: 9 (x, y) landmark coordinates normalized
to [0, 1] in the letterboxed square input space.
train.py implements:
- Aspect-ratio-preserving preprocessing — images are padded to square on
both sides of the short edge (letterbox) instead of being stretched, so the
cat's facial anatomy is never distorted. Inference (
main.py) uses the same letterboxing and maps predictions back to original image coordinates. - Augmentations — random horizontal and vertical flips (with correct left/right relabeling of the symmetric eye/ear landmarks) and random rotation in all directions (−180°…180°, canvas expanded so nothing is cut off).
- Random crop samples — each source image also contributes a randomly cropped version whose crop edges always stay outside the annotated landmark bounding box plus a margin, so facial points are never lost.
uv sync --group train
python train.py --data-dirs data/CAT_00 data/CAT_01 ... --variant both --epochs 30--variant small|large|both selects which model(s) to train. Best-validation
weights are checkpointed to cat_landmark_{variant}.pth and exported to
cat_landmark_{variant}.onnx.
python main.py test.jpg --variant small # or --variant large
python main.py photo.jpg --model path/to/model.onnx --output annotated.jpgOutputs landmark coordinates to stdout and saves an annotated image.
onnxruntimeopencv-python- training additionally needs
torch,torchvision,onnx(see thetraindependency group)
Install with uv:
uv sync