shima78/Face-Recognition

Face detection, tracking, re-identification, and open-set recognition (SPL/MPL)

★ 0Forks 0PythonGitHub ↗Compare

README

Face Detection, Tracking, Re-Identification & Open-Set Recognition

A small face-recognition pipeline plus a separate tabular open-set recognition (OSR) challenge.

Part 1 — Face pipeline (webcam / video)

  • face_detector.py — FaceDetector: MTCNN-based face detection, landmark alignment to a fixed square size, and frame-to-frame tracking via OpenCV template matching (cv2.matchTemplate) with automatic re-detection when the match score drops below tm_threshold.
  • face_recognition.py / my_face_recognition.py — two versions of the same models (the latter uses package-relative imports, i.e. is meant to be run as part of the cvproj_exc package from src/):
    • FaceNet — wraps an ONNX ResNet50 face-embedding model (data/resnet50_128.onnx), producing L2-normalized 128D embeddings.
    • FaceRecognizer — supervised identification via kNN over stored (label, embedding) pairs, fusing predictions from color and grayscale embeddings, with open-set rejection ("unknown") when distance/probability thresholds (max_distance, min_prob) aren't met.
    • FaceClustering — unsupervised k-means over embeddings for re-identification without labels, with open-set rejection via a distance threshold tau, k-means restarts (fit_minimum_objective), and convergence plotting.
  • classifier.py — NearestNeighborClassifier, a thin wrapper around cv2.ml.KNearest used by the DIR-curve evaluation.
  • training.py / train_all.py — capture frames from a webcam or an image-sequence "video", track the face, and incrementally fit either the identification gallery (--mode ident) or the clustering gallery (--mode cluster); saves the trained model on exit.
  • test.py / test_tracking.py — same capture/track loop as training, but overlays live identification or clustering predictions on the video feed instead of updating the models.
  • cam_test.py — minimal OpenCV webcam smoke test (no detection/recognition).
  • cma_sharing — local note: usbipd command to attach a USB webcam to WSL for the scripts above.

Evaluation / analysis

  • evaluation.py — OpenSetEvaluation: computes a DIR curve (identification rate vs. false alarm rate) by sweeping similarity thresholds, and prints the best operating points for FAR ≤ 1% and IR ≥ 90%.
  • dir_curve.py — runs OpenSetEvaluation on data/evaluation_{train,test}_data.pkl and plots the DIR curve.
  • find_eval_tau.py — analyzes known-vs-unknown embedding distance distributions on the evaluation data (histograms, tau sweep, PCA visualization) to help pick a good rejection threshold.
  • plot_clusters.py — trains FaceClustering on data/train_data and evaluates known/unknown separation on data/test_data, with the same histogram/tau-sweep/PCA diagnostics.
  • reidentification_test.py — end-to-end re-identification benchmark: enrolls people from data/train_data into FaceClustering, builds a cluster→person mapping by majority vote, then evaluates rank-1 accuracy and unknown-detection rate on data/test_data (including confusion summaries).

Part 2 — Open-set recognition on tabular features (challenge_train_data.csv)

A 128D-feature, multi-class dataset with a dedicated "unknown" label (-1) for known-unknown classes (KUCs), used to benchmark two open-set strategies:

  • osr_learning.py
    • SPLClassifier (Single Pseudo Label) — trains one RandomForestClassifier over known classes plus a single merged "unknown" pseudo-class; confidence combines known-vs-unknown probability gap and prediction entropy.
    • MPLClassifier (Multi Pseudo Label) — trains one LogisticRegression where every known-unknown sample gets its own unique pseudo-label (so the "unknown" region is modeled as many small classes rather than one), then maps any negative-label prediction back to -1.
    • spl_training / mpl_training — the required benchmark interface: given (x_train, y_train), return a predict_fn(x_test) -> (y_pred, y_score).
    • load_challenge_train_data() — loads challenge_train_data.csv (last column = label).
  • tune_osr.py — an independent, lighter-weight prototype-based baseline (cosine similarity to per-class/pseudo-class prototypes) used to sweep hyperparameters (tau, number of pseudo-clusters for MPL) and pick operating points for target false-alarm rates.
  • eval_osr.py — trains spl_training/mpl_training on a held-out split of the challenge data and reports AUC (known vs. unknown), balanced rank-1 accuracy, and DIR@FAR at several operating points.
  • test_osr_learning.py — unit tests asserting spl_training/mpl_training keep the required interface (callable, returns (y_pred, y_score) as 1D arrays of the right length) and correctly predict at least some UNKNOWN_LABEL.

Data layout expected by Config (config.py)

Config.PROJECT_DIR resolves two levels above config.py (i.e. exercise-04-data/), with:

exercise-04-data/
├── data/
│   ├── train_data/<person>/*.jpg        # enrollment images per identity
│   ├── test_data/<person>/*.jpg         # test images (some identities unseen in train)
│   ├── resnet50_128.onnx                # FaceNet embedding model
│   ├── clustering_gallery.pkl           # saved FaceClustering state
│   ├── recognition_gallery.pkl          # saved FaceRecognizer state
│   ├── evaluation_train_data.pkl        # (embeddings, labels) for DIR-curve eval
│   ├── evaluation_test_data.pkl
│   └── challenge_train_data.csv
└── src/cvproj_exc/                      # this repo

challenge_train_data.csv is gitignored and expected to sit alongside these scripts locally (osr_learning.py's load_challenge_train_data() reads it via Config.CHAL_TRAIN_DATA, and tune_osr.py takes a --csv path to it directly). The data/ directory itself (images, .pkl/.onnx model files) is not part of this repo.

Usage

pip install -r requirements.txt

# Enroll a person via webcam, then test recognition
python training.py --mode ident --video none --label Alice
python test.py --mode ident --video none

# Unsupervised clustering / re-identification
python training.py --mode cluster --video none
python reidentification_test.py

# Open-set evaluation (needs data/evaluation_{train,test}_data.pkl)
python dir_curve.py
python find_eval_tau.py

# Tabular open-set challenge
python osr_learning.py                 # sanity-run SPL + MPL on dummy test data
python eval_osr.py                     # held-out split benchmark
python tune_osr.py --csv challenge_train_data.csv --method mpl
python -m unittest test_osr_learning.py

Requirements

See requirements.txt: matplotlib, mtcnn[tensorflow], opencv-python, pandas, scikit-learn.

Contributors

shima78

Issues