Hurst exponent less than 0 or greater than 1?

#7 · closed · 7 comments

View on GitHub ↗

blu3r4y

I noticed that the `compute_Hc` result for some time series is less than 0 or greater than 1. Is this behavior intended? MWE to reproduce this bug: ``` import hurst import numpy as np print("numpy", np.__version__) print("hurst", hurst.__version__) print() np.random.seed(988) H, _, _ = hurst.compute_Hc(np.random.uniform(size=100), kind="random_walk", simplified=True) print(H) # -0.017687382184009826 np.random.seed(916) H, _, _ = hurst.compute_Hc(np.random.uniform(size=100), kind="random_walk", simplified=False) print(H) # -0.011722357538317393 np.random.seed(164) H, _, _ = hurst.compute_Hc(np.random.exponential(1, size=100), kind="change", simplified=True) print(H) # 1.0118591069505447 ```

Comments

Mottl

Hi, `np.random.uniform` doesn't produce random walk and it also doesn't produce changes since `np.random.uniform` by default generates values in [0, 1) range but changes should contain both positives and negatives ones. For the same reason you can't use `np.random.exponential` since it generates only positive values.

blu3r4y

Okay, so I do understand that change can only be used with series that also contain negative values, right? However, I am still confused by the random walk kind. What assumption exactly must hold for a time series to be able to calculate the hurst exponent? In other words, why would the following not be considered a random walk? ``` np.random.seed(988) plt.plot(np.random.uniform(size=100)) ``` ![rnd-walk](https://user-images.githubusercontent.com/10400532/66637804-0fcb6600-ec14-11e9-9df6-9b3af0e4e1aa.png)

Mottl

A random walk is a cumulative sum of random values. Refer to README: https://github.com/Mottl/hurst#kinds-of-series: `np.cumsum(np.random.randn(...))`

blu3r4y

Okay, so here are two random walks that are out of range. Maybe this is due to some numerical problems? ``` np.random.seed(26240) x1 = np.cumsum(np.random.randn(100)) np.random.seed(81984) x2 = np.cumsum(np.random.randn(100)) H1, _, _ = hurst.compute_Hc(x1, kind="random_walk", simplified=False) H2, _, _ = hurst.compute_Hc(x2, kind="random_walk", simplified=True) print(H1) # 1.018265313908482 print(H2) # 1.008202895027869 ``` ![rnd-walk](https://user-images.githubusercontent.com/10400532/66639228-c16b9680-ec16-11e9-9d18-febedd623340.png)

Mottl

Both blue and orange lines look to me as persistent. So H is about 1. The more points you generate the closer H will be to 0.5.

blu3r4y

So this might be a numerical problem of the implementation and H should just be clamped to 1 in these scenarios?

Mottl

_H_ is just a slope of a linear regression. The more time-series observations you have the longer interval _n_ you have and eventually the better estimate of _H_ you can get: ![](https://wikimedia.org/api/rest_v1/media/math/render/svg/6e6b162d820f98b5a3b42be6c64beff9e717e0c5) Whether you want to clip _H_ to (0, 1) range or not depends on your task. If you doubt the correctness of the calculations you can check https://en.wikipedia.org/wiki/Hurst_exponent#Rescaled_range_(R/S)_analysis and compare with the actual Python code.