This repository contains the solution for the FlamApp Research & Development assignment. This project reverse-engineers a complex parametric curve by extracting three hidden parameters (xy_data.csv). By deploying a spatial KD-Tree alongside a Differential Evolution global optimizer, this pipeline bypasses data-shuffling traps and minimizes the
The assignment provided a dataset of 1,500 coordinates and required an algorithm to:
- Extract the true values of
$\theta$ ,$M$ , and$X$ from a given set of complex parametric equations. - Ensure
$\theta$ is strictly bounded between$0^\circ < \theta < 50^\circ$ . - Evaluate the accuracy of the extracted parameters strictly using
$L_1$ (Manhattan) distance. - Provide the mathematical justification for the estimation logic used.
graph LR
subgraph Data Pipeline
A[Raw Dataset xy_data.csv] -->|Shuffled Data| B(Spatial Partitioning scipy.spatial.cKDTree)
B --> C[(Target Point Cloud)]
end
subgraph Optimization Loop
D((Differential Evolution)) -->|Candidate Params θ, M, X| E[Parametric Curve Generator]
E --> F([Predicted Point Cloud])
F --> G{L1 Distance Manhattan Norm}
C -.->|Nearest Neighbor Query| G
G -->|Error Penalty| D
end
D ===>|Convergence| H((Optimal Parameters θ=30, M=0.03, X=55))
%% Styling
classDef data fill:#f9f6f0,stroke:#333,stroke-width:2px,color:#000;
classDef engine fill:#e1f5fe,stroke:#0288d1,stroke-width:2px,color:#000;
classDef result fill:#e8f5e9,stroke:#388e3c,stroke-width:3px,color:#000;
class A,B,C data;
class D,E,F,G engine;
class H result;
The provided parametric equations were translated into a modular Python script using numpy for vectorized mathematical operations.
-
Unit Conversion: The parameter
$\theta$ was strictly bounded in degrees ($0^\circ < \theta < 50^\circ$ ). The script dynamically converts this to radians during computation to ensure correct trigonometric mapping. -
Spatial Distance Calculation: The target CSV dataset rows were heavily randomized (shuffled
$X$ and$Y$ values). Instead of attempting flawed textbook sorting on a parametric loop, I implemented a KD-Tree (scipy.spatial.cKDTree) to treat the target data as a spatial point cloud. The custom$L_1$ loss function queries this tree to find the true nearest-neighbor distance for every predicted point, making the pipeline completely immune to data shuffling.
Because the parametric equations involve complex exponential and sinusoidal combinations, standard local optimizers (like gradient descent) are highly susceptible to getting trapped in local minima. To guarantee an optimal fit, I implemented SciPy's Differential Evolution algorithm.
- Why Differential Evolution? It is a robust, stochastic global optimizer that efficiently searches large, non-convex mathematical spaces.
- By feeding it the strict parameter boundaries provided in the assignment, the algorithm successfully converged on the global minimum.
In standard regression problems, datasets are sequentially ordered. To calculate the error of a prediction array (xy_data.csv was highly shuffled, comparing index
To solve this, I modeled the target dataset as a rigid 2-dimensional point cloud (cKDTree constructs a binary space-partitioning tree:
- It splits the space along the
$x$ -axis at the median point. - It recursively splits the sub-spaces along the
$y$ -axis, alternating axes until every coordinate is mapped to a leaf node.
This spatial index drops the time complexity of searching for the nearest neighbor from
For every point generated by the parametric guess
The assignment strictly required the
By bypassing array index constraints, the global optimizer minimizes the true spatial error across all
The final mathematical objective function the algorithm minimizes is:
This custom spatial mapping successfully neutralized the randomized dataset, driving the mathematical error to the exact global minimum.
After implementing the spatial KD-Tree, the global optimizer successfully converged with a near-zero final
The optimal extracted parameters directly from the terminal output are:
-
Theta (
$\theta$ ): 29.9997 degrees - M: 0.0300
- X: 54.9986
(Note: These values mathematically converge on the exact parameters of θ = 30°, M = 0.03, and X = 55.)
The plot below demonstrates the success of the spatial KD-Tree pipeline. The algorithm successfully ignored the randomized row indices and converged on the true global shape of the parametric curve.
The red line represents the dynamic output of the Differential Evolution optimizer, cutting precisely through the raw, shuffled target data.
Desmos Link : click-here
Below is the exact plain-text string generated by the pipeline with the optimized variables plugged in. Paste this into Desmos with the bounds set to 6 <= t <= 60 to visualize the exact curve overlap.
(t * cos(0.5236) - e^(0.0300 * |t|) * sin(0.3 * t) * sin(0.5236) + 54.9986, 42 + t * sin(0.5236) + e^(0.0300 * |t|) * sin(0.3 * t) * cos(0.5236))
- Clone this repository and enter the directory:
git clone https://github.com/DevaRajan8/flam-assignment.git cd flam-assignment - Install the required dependencies:
pip install -r requirements.txt
- Add the dataset: Ensure the official
xy_data.csvis placed in the root directory (Ignored via.gitignoreto maintain dataset privacy). - Execute the pipeline: Run the following command to execute the solver, calculate the parameters, and generate the final overlay plot:
python src/optimizer.py



