Release grumpy 0.1.1

#10 · closed · 5 comments

View on GitHub ↗

Bisaloo

First release: * [x] `usethis::use_cran_comments()` * [x] Update (aspirational) install instructions in README * [x] Proofread `Title:` and `Description:` * [x] Check that all exported functions have `@return` and `@examples` * [x] Check that `Authors@R:` includes a copyright holder (role 'cph') * [x] Check [licensing of included files](https://r-pkgs.org/license.html#sec-code-you-bundle) * [x] Review <https://github.com/DavisVaughan/extrachecks> Prepare for release: * [x] `git pull` * [x] `urlchecker::url_check()` * [x] `devtools::build_readme()` * [x] `devtools::check(remote = TRUE, manual = TRUE)` * [x] `devtools::check_win_devel()` * [x] `git push` Submit to CRAN: * [x] `usethis::use_version('patch')` * [x] `devtools::submit_cran()` * [x] Approve email Wait for CRAN... * [x] Accepted :tada: * [x] `usethis::use_github_release()` * [x] `usethis::use_dev_version(push = TRUE)`

Comments

eddelbuettel

Hi @Bisaloo -- congrats on landing this on CRAN. Are you aware of the RcppCNPy package doing just that based on older external C library (`cnpy`) it wraps? You may well have more complete support here now---that was just something I once needed to pick up data from Python (and years before we had `reticulate`, of course) which also does that for us.

Bisaloo

Hi, yes, thanks for reaching out. It might have been better to communicate with you earlier. We tried RcppCNPy and got entangled in https://github.com/eddelbuettel/rcppcnpy/issues/21. The problem was that the npy files were written by an external python library so it was difficult to save it with a different bytesize. Especially because they were using lower-precision floats by design because higher-precision was not needed for their use case and they wanted to save disk space. I wrote quickly a very simple proof of concept of how we could read our float32 npy objects in R, with `readBin()` and realized there was a lot of duplication with our existing [Rarr bioconductor package](https://bioconductor.org/packages/release/bioc/html/Rarr.html). So this justified spending a bit more time to migrate some code from Rarr and pushing it to CRAN as I can now trim and greatly simplify Rarr (https://github.com/Huber-group-EMBL/Rarr/pull/174). I also saw that the underlying cnpy library was not actively developed and it wasn't clear if they were just happy with the current state, or if it was no longer maintained. Please let me know if I'm missing something. I would be happy to mention RccpCNPy as an alternative somewhere but I wanted to sync with you first to not misrepresent what RcppCNPy can & can't do. This has happened to me in the past when people published alternatives to my existing packages and it's frustrating.

Bisaloo

Oh, and one extra thing is that for Rarr at least, we definitely need more than 3 dimensions. grumpy supports an arbitrary number of dimensions.

eddelbuettel

All good -- RcppCnpy was made for a world where I needed (simple) n by k matrices from a Python NumPy process, and `cnpy` fit that bill. As you can from the issues, there have always been some corner cases from deeper / further Python use and who got to cover those by the 80/20 rule. And ... for most basic cases `reticulate` also helps. Which leads to my next curious question: why did you not just rely on it? Same story, didn't help enough with `zarr` data and `rarr`?

Bisaloo

To be honest, part of the reason is: the format is so simple that we don't have to bother with reticulate, so why should we? We also had most of the code already written for Rarr before noticing it can be re-used for npy files. I have also mentioned some reasons in the [design vignette](https://hugogruson.fr/grumpy/articles/design.html#why-not-use-reticulate). I think we can also get better performance (reading speed + peak memory consumption) with suggestions such as #11.