DevaRajan8/flam-assignment

★ 0Forks 0PythonGitHub ↗Compare

README

FlamApp AI - R&D Assignment

Python 3.12 SciPy uv CI

Project Overview

This repository contains the solution for the FlamApp Research & Development assignment. This project reverse-engineers a complex parametric curve by extracting three hidden parameters ($\theta$, $M$, $X$) from a raw, shuffled dataset (xy_data.csv). By deploying a spatial KD-Tree alongside a Differential Evolution global optimizer, this pipeline bypasses data-shuffling traps and minimizes the $L_1$ distance to a near-zero convergence.

The Challenge

The assignment provided a dataset of 1,500 coordinates and required an algorithm to:

  • Extract the true values of $\theta$, $M$, and $X$ from a given set of complex parametric equations.
  • Ensure $\theta$ is strictly bounded between $0^\circ < \theta < 50^\circ$.
  • Evaluate the accuracy of the extracted parameters strictly using $L_1$ (Manhattan) distance.
  • Provide the mathematical justification for the estimation logic used.

Methodology & Technical Approach

Pipeline Architecture

graph LR
    subgraph Data Pipeline
        A[Raw Dataset xy_data.csv] -->|Shuffled Data| B(Spatial Partitioning scipy.spatial.cKDTree)
        B --> C[(Target Point Cloud)]
    end

    subgraph Optimization Loop
        D((Differential Evolution)) -->|Candidate Params θ, M, X| E[Parametric Curve Generator]
        E --> F([Predicted Point Cloud])
        F --> G{L1 Distance Manhattan Norm}
        C -.->|Nearest Neighbor Query| G
        G -->|Error Penalty| D
    end

    D ===>|Convergence| H((Optimal Parameters θ=30, M=0.03, X=55))

    %% Styling
    classDef data fill:#f9f6f0,stroke:#333,stroke-width:2px,color:#000;
    classDef engine fill:#e1f5fe,stroke:#0288d1,stroke-width:2px,color:#000;
    classDef result fill:#e8f5e9,stroke:#388e3c,stroke-width:3px,color:#000;

    class A,B,C data;
    class D,E,F,G engine;
    class H result;

Loading

alt text

1. The Mathematical Engine (math_utils.py)

The provided parametric equations were translated into a modular Python script using numpy for vectorized mathematical operations.

  • Unit Conversion: The parameter $\theta$ was strictly bounded in degrees ($0^\circ < \theta < 50^\circ$). The script dynamically converts this to radians during computation to ensure correct trigonometric mapping.
  • Spatial Distance Calculation: The target CSV dataset rows were heavily randomized (shuffled $X$ and $Y$ values). Instead of attempting flawed textbook sorting on a parametric loop, I implemented a KD-Tree (scipy.spatial.cKDTree) to treat the target data as a spatial point cloud. The custom $L_1$ loss function queries this tree to find the true nearest-neighbor distance for every predicted point, making the pipeline completely immune to data shuffling.

2. The Optimization Engine (optimizer.py)

Because the parametric equations involve complex exponential and sinusoidal combinations, standard local optimizers (like gradient descent) are highly susceptible to getting trapped in local minima. To guarantee an optimal fit, I implemented SciPy's Differential Evolution algorithm.

  • Why Differential Evolution? It is a robust, stochastic global optimizer that efficiently searches large, non-convex mathematical spaces.
  • By feeding it the strict parameter boundaries provided in the assignment, the algorithm successfully converged on the global minimum.

Deep Dive: The Mathematics of the KD-Tree Solution

The Flaw of Standard Array Loss

In standard regression problems, datasets are sequentially ordered. To calculate the error of a prediction array ($P$) against a target array ($T$), one would simply calculate the column-wise residual: $$Error = \frac{1}{N} \sum_{i=1}^{N} |P_i - T_i|$$ However, because the xy_data.csv was highly shuffled, comparing index $i$ of the prediction to index $i$ of the target results in catastrophic geometric distortion, especially for a parametric curve that loops backward upon itself.

The Spatial Partitioning ($k$-d tree)

To solve this, I modeled the target dataset as a rigid 2-dimensional point cloud ($k=2$). The cKDTree constructs a binary space-partitioning tree:

  1. It splits the space along the $x$-axis at the median point.
  2. It recursively splits the sub-spaces along the $y$-axis, alternating axes until every coordinate is mapped to a leaf node.

This spatial index drops the time complexity of searching for the nearest neighbor from $O(N)$ to $O(\log N)$, allowing the differential evolution optimizer to evaluate thousands of candidate curves per second.

The $L_1$ Distance Metric (Manhattan Norm)

For every point generated by the parametric guess $P_{pred} = (x_{pred}, y_{pred})$, the tree is queried for the nearest target coordinate $P_{target} = (x_{target}, y_{target})$.

The assignment strictly required the $L_1$ distance. The distance between the predicted point and the physically closest target point is calculated using the Manhattan norm: $$D_{L_1}(P_{pred}, P_{target}) = |x_{pred} - x_{target}| + |y_{pred} - y_{target}|$$

The Final Objective Function

By bypassing array index constraints, the global optimizer minimizes the true spatial error across all $N$ generated points.

The final mathematical objective function the algorithm minimizes is:

$$ \mathcal{L}_{total} = \frac{1}{N} \sum_{i=1}^{N} \min_{j \in \text{Target Cloud}} \left( |x_{pred}^{(i)} - x_{target}^{(j)}| + |y_{pred}^{(i)} - y_{target}^{(j)}| \right) $$

This custom spatial mapping successfully neutralized the randomized dataset, driving the mathematical error to the exact global minimum.


Final Results and Plots

Desmos Plot

alt text

After implementing the spatial KD-Tree, the global optimizer successfully converged with a near-zero final $L_1$ Error of 0.026726.

The optimal extracted parameters directly from the terminal output are:

  • Theta ($\theta$): 29.9997 degrees
  • M: 0.0300
  • X: 54.9986

alt text

(Note: These values mathematically converge on the exact parameters of θ = 30°, M = 0.03, and X = 55.)

Visual Proof of Convergence

The plot below demonstrates the success of the spatial KD-Tree pipeline. The algorithm successfully ignored the randomized row indices and converged on the true global shape of the parametric curve.

alt text

The red line represents the dynamic output of the Differential Evolution optimizer, cutting precisely through the raw, shuffled target data.

Desmos Verification String

Desmos Link : click-here

Below is the exact plain-text string generated by the pipeline with the optimized variables plugged in. Paste this into Desmos with the bounds set to 6 <= t <= 60 to visualize the exact curve overlap.

(t * cos(0.5236) - e^(0.0300 * |t|) * sin(0.3 * t) * sin(0.5236) + 54.9986, 42 + t * sin(0.5236) + e^(0.0300 * |t|) * sin(0.3 * t) * cos(0.5236))

How to Run Locally

  1. Clone this repository and enter the directory:
    git clone https://github.com/DevaRajan8/flam-assignment.git
    cd flam-assignment
  2. Install the required dependencies:
    pip install -r requirements.txt
  3. Add the dataset: Ensure the official xy_data.csv is placed in the root directory (Ignored via .gitignore to maintain dataset privacy).
  4. Execute the pipeline: Run the following command to execute the solver, calculate the parameters, and generate the final overlay plot:
    python src/optimizer.py

Contributors

DevaRajan8

Issues