Knowledge, as justified true belief, is normally seen as resulting from lucky guesses, but optimality provides a sounder basis to epistemology and, therefore, this can be the standard against which the inherent dimensionality of data may be compared Information theory and dimensionality of space, Kak (2020)
This repository is a collection of observations, codes and references from a personal approach to data science from the persective of Information Theory. The process to develop these notes was non-linear, often following a chain of somewhat connected topics related to Information Theory. Many of the techniques, codes and approaches that we employed can be used across different problems. That was the motivation to attempt to write a "library" of functions. From exploratory data analysis (EDA), feature engineering, model training, model selection, evaluation, pipeline and story telling, among others, we will explore connections with concepts and practical aspects of Information Theory. We will explore a few applications in network theory, flowck dynamics, reinforcement learning etc. Real world applications will be highlighted.
By considering the event as primary, we can measure its information as related to its frequency: the less likely the event, the information is greater. The logarithm measure associated with information has become widely accepted because it is additive. In other words, additivity of information is the underlying unstated assumption in the information-theoretic approach to structure and since information is related to the experimenter, this constitutes a subject-centered approach to reality. The entropy, or the average information associated with a source, is maximized when the events have the same probability. Information theory and dimensionality of space, Kak (2020)
We will use a set of well-known benchmark datasets (e.g Titanic data set from Kaggle) to build our use cases and will develop a library of functions in python to streamline the approach. We will write the librabry following a basic form of Test driven development (TDD) and pytest. Additionally, a wiki with in-depth explanation of concepts, tools and algorithms will develop some concepts and details further.
Thisalgorithmbelongsideologicallytothatphilosophicalschool that allows wisdom to emerge rather than trying to impose it, that emulates nature rather than trying to control it, and that seeks to make things simpler rather than more complex. Once again nature has provided us with a techniquefor processing informationthat is at once elegant and versatile. Particle Swarm Optimization, Kennedy and Eberhart (1995)
Can we define a quantity which will measure, in some sense, how much information is “produced” [by such a process], or better, at what rate information is produced? A Mathematical Theory of Communication, Shannon (1948)
Data Science leverages models to convert raw data into (actionable) information. How much information is encoded in the data at hand is a great starting question. In this section we introduce the concept of entropy in the context of information theory. Then we discuss the challenges with continuous and mixed distributions and implement a method to estimate entropy and mutual information for mixed distributions. We consider the relevance matrices, a matrix buitl from the pair-wise mutual information from a data set and explore the behavior for different MI thresholds link. The animation below shows the change of the relevance matrix for the CTR dataset as we vary the mutual information threshold more
In physics, one definition of degrees of freedom for a general mechanical system is the following:
The number of degrees of freedom of a system is the number of independent variables that must be specified to define completely the condition of the system Reference0.
A similar concept exists in data science - given a data set and a problem (context), what's the minimum number of independent variables needed to specify completely the system so that we can explain/predict the outcome of a future (unseen) state.
Often times, we are dealing with a data set with several features, but only a few convey relevant information about the problem in hand. Ideally, we would like to identify the number of degrees of freedom of our system, from the data set. The dimension of the configuration space makes the concept of number of degrees of freedom more precise with the following intuitive notions Information Dimension and the Probabilistic Structure of Chaos, Farmer (1982).
- Direction, related to the topological dimension
- Capacity, related to fractal dimension
- Measure, related to information dimension.
We might be pardoned for supposing that in a space of infinite dimension we should find the Absolute and Unconditioned if anywhere, but we have reached an opposite conclusion. This is the most curious thing I know of in the Wonderland of Higher Space. Properties of the locus r=constant in space of n dimensions, Heyl (1897)
Our intution of dimensions is heavely influenced by our 3-dimensional world. We tend to extrapolate what happens in 1,2 and 3 dimension to higher dimensions and often our intuition is wrong. In this section, we will consider the volume of the unit ball in n dimensions, the ratio of this volume vs the volume of the n dimensional cube that contains it. As we will see, this ratio tends to zero as the dimension increases - most of the volume of the n-cube is located outside the n-ball, as the cube expands the ball shrinks, both at impressive rates. Reference1
There is little difficulty in applying the calculus of variations, dynamic programming, or any of a number of direct methods, to systems described by state vectors of low dimension to obtain efficient computational techniques. When, however, the dimension is “large” (a relative rather than absolute property), or infinite, the “curse of dimensionality” prevents the direct use of general methods. A new type of approximation leading to reduction of dimensionality in control processes, Bellman (1969
The behaviour of points/data in high dimensions has a profound impact in data analysis. Many data analysis and machine learning algorithm rely on the concept of neighbors - points within a certain distance. If we consider the n-ball around a fixed point, as n increases the chances that a another point will land within the n-ball (i.e. become neighbors) decreases very fast. The whole notion of neighborhood needs to be revisited.
In this repository we will discuss the volume and area of an n-ball both analytically and with a Monte Carlo simulation. Below, we see a simulation of the MC computation of the ratio of the volume n-ball vs n-cube for different values of p-norm more
Noninteger dimensions provide explanation for recursion and scale-invariance in complex systems—biological, physical, or engineered—and explain fractal behavior, examples of which include the Mandelbrot set, patterns in Romanesco broccoli, snowflakes, the Nautilus shell, complex computer networks, brain structure, as well as the filament structure and distribution of matter in cosmology. Information theory and dimensionality of space, Kak (2020)
In this section we will construct an example of a fractal: the Henon Map and we will compute its dimension. Below is an animation of the Henon Map for n*1000 iterations (n=1,1000) more.

The link between fractals and information theory is the notion of information dimension:
The information dimension, D1, measures the fractal dimension of a probability distribution, and relates the growth of Shannon entropy to how the system under study is discretized A new fractal index to classify forest disturbance and anthropogenic change
The direct physical relevance of the information dimension is in measurement. Knowledge of the information dimension of an attractor allows an observer to estiamte the information gaines when a measurement is made at a given level of precision Information Dimension and the Probabilistic Structure of Chaos, Farmer (1982)
It does not seem a too-large leap of logic to suppose that some same rules underlie animal social behavior,includingherds,schools,andflocks,andthatofhumans. AssociobiologistE.0.Wilson[9] has written, in reference to fish schooling, “In theory at least, individual members of the school can profit from the discoveries and previous experience of al other members of the school during the searchforfood. Thisadvantagecanbecome decisive,outweighingthedisadvantagesof competition for food items, whenever the resource is unpedictably distributed in patches” (p.209). This statement suggeststhat social sharingof informationamong conspeciatesoffers an evolutionaryadvantage:this hypothesiswas fundamentalto the developnmt of particle swarm optimization Particle Swarm Optimization, Kennedy and Eberhart (1995).
In the animation below, we can see the particle swarm optimization algorithm at work, finidn the minimum point in a non-linear, continuous 2D surface. WE can see how the particles start in a random configuration and as they collect the particle best and share them to find the global best after a few iterations the swarm lands very close to the optimal value.
Physical space of course affects informational inputs, but it is arguably a trivial component of psychological experience. Humans learn to avoid physical collision by an early age, navigation of n-dimensional psychosocial space requires and many of us never seem to acquire quite al the skills we need Particle Swarm Optimization, Kennedy and Eberhart (1995)
Holland’s chapter on the “optimum allocation of trials” [5] reveals the delicate balance between conservative testing of known regions versus risky exploration of the unknown. It appears that the current version of the paradigm allocates trials nearly optimally. The stochastic factors allow thorough search of spaces between regions that have been found to be relatively good, and the momentum effect caused by nmhfying the extant velocities rather than replacing them results in overshooting,or exploration of unknown regions of the problem domain. Particle Swarm Optimization, Kennedy and Eberhart (1995)
Ant Colony Optimization algorithms, that is, instance of the ACO metaheuritics (...) use a population of ants to collectively solve the optimization problem under consideration (...). Information collected by the ants during the search process is stored in pheromone trails
$\tau_{i,j}$ associated to connections$l_{i,j}$ . Pheromone trails encode a long-term memory about the whole ant search process. Anto Colony Optimization: A new Metaheuristic Dorigo and Dicaro (1999)
The animation below, shows a toy model of the problem, with two nodes and two paths of different lengths. The ants will choose a path based on a probability distribution propotional to the value of
Q-learning (Watkins, 1989) is a simple way for agents to learn how to act optimally in controlled Markovian domains. It amounts to an incremental method for dynamic programming which imposes limited computational demands. It works by successively improving its evaluations of the quality of particular actions at particular states Q-Learning, Watkins and Dayan, 1992
Suppose we need to find the location of the next store for a national retail chain withthe condition of being located in the largest area without presence; or suppose we want to locate a waste facility that is as far as possible from current houses. The Largest Empty Circle answers those questions.
The largest empty circle (LEC) problem is defined on a set P and consists of finding the largest circle that contains no points in P and is also centered inside the convex hull of PThe Largest Empty Circle Problem, Schuster (2008).
In order to facilitate the visualization, we'd like to locate random points that are not clustered or very close to each other. We use a Possion Disk distribution to simulate a blue noise sample pattern.
Blue noise sample patterns—for example produced by Poisson disk distributions, where all samples are at least distance r apart for some user-supplied density parameter r—are generally considered ideal for many applications in rendering. Fast Poisson Disk Sampling in Arbitrary Dimensions, Bridson (2007)
[For those channels] by analogy to the combinatorial problem of construction optimal codes capable of correcting s reversals, we will consider the problem of construction optimal codes capable of correcting deletions, insertions, and reversals. Binary Codes Capable of Correcting Deletions, Insertions and Reversals. Leveinshten (1966)
In this example we will consider a corpus of words in the english language and compute the Levenshtein distance to build word paths from a source word to a target words. Each move from word to word corresponds to a insertion a delation or a reversal (exchange). Below we see a representation of the word network revealing a central cluster of connected words and an outer ring of disconnected words.
The animation shows a path between the word well and the word john. For this particular case, we see only exhanges since all words have the same length.

The problem of multidimensional scaling, broadly stated, is to find n points whose interpoint distances match in some sense the experimental dissimilarities of n objects. Instead of dissimilarities the experimental measurements may be similarities, confusion probabilities, interaction rates between groups, correlation coefficients, or other measures of proximity or dissociation of the most diverse kind. Whether a large value implies closeness or its opposite is a detail and has no essential significance. What is essential is that we desire a monotone relationship, either ascending or descending, between the experimental measurements and distances in the configuration MULTIDIMENSIONAL SCALING BY OPTIMIZING GOODNESS OF FIT TO A NONMETRIC HYPOTHESIS, Kruskal (1964)
We have used a multidimensional scaling to represent the words from the entire corpus and the path from source to target.

We employ maximum likelihood estimation (MLE) to find the parameters that best fit a model built from first principles that describes the probability of purchase for a Recency and Frequency model. The model is a combination of a beta-gemoetric and negative binomial distributions, with exact and analitically explicit equations. The model is highly interpretable with acceptable accuracy (~2% total forecast) wiki notebook
Despite almost 10 billion ad impressions per day in our dataset, hundreds of millions of unique user ids, millions of unique pages, and millions of unique ads, combined with the lack of easily generalizable features, makes sparsity a significant problem Simple and Scalable Response Prediction for Display Advertising, Chapelle et al, (2014)







