Comments (3)
Hi Astrid! Thanks for using litstudy and thanks for reporting this issue!
Unfortunately, at the moment there is no functionality to see which papers were removed when taking the union of multiple document sets.
Issue #68 discussed a similar problem where the is now way to find the papers removed by unique()
. An idea there was to add a duplicates()
method that returns the papers removed by unique()
(such that len(docset) == len(docset.unique()) + len(docset.duplicates())
. Something similar could be implemented for union()
.
We are open to contributes and will accept relevant pull requests that add this functionality.
from litstudy.
Good to know, thanks! However, I'm actually more interested in the documents that are kept after the union (so not the removed duplicates); e.g. to know which documents I should look into for my review, and thus also the titles of the documents that the different kinds of histograms are based on. Is that possible to do with LitStudy?
from litstudy.
You can always print the documents like this:
docs_csv = docs_ieee | docs_springer
for doc in docs_csv:
print(doc.title)
Would that work? Each document has many attribute that you can access (such as the title, authors, publisher, etc.). See here: https://nlesc.github.io/litstudy/api/types.html#litstudy.types.Document
from litstudy.
Related Issues (20)
- 'No Edges Given' for Network Analysis HOT 4
- ValueError: n_components must be < n_features; got 50 >= 47 HOT 2
- `build_corpus` always removes words having a frequency below 5 HOT 4
- Different results from unique() and difference of deduplicated set HOT 2
- module 'networkx' has no attribute 'to_scipy_sparse_matrix' HOT 2
- Incompability with gensim 4 HOT 1
- Unexpected results from litstudy.plot_author_histogram() HOT 2
- Support for google scholar HOT 1
- refine_scopus - low it/s speed; necessary to refine every time? HOT 1
- TypeError: object of type 'method' has no len() HOT 1
- Saving language models
- Documentation on search_ function queries HOT 3
- Search_semanticscholar with list
- Scopus400Error: Error translating query - Refining results with "source title" query argument HOT 6
- train_lda_model() fails to access gensim HOT 3
- Scopus400Error: Exceeds the maximum number allowed for the service level. HOT 1
- Scopus exceeds csv field limit
- SemanticScholar search optimization HOT 2
- DocumentIdentifier.matches() is case-sensitive
Recommend Projects
-
React
A declarative, efficient, and flexible JavaScript library for building user interfaces.
-
Vue.js
🖖 Vue.js is a progressive, incrementally-adoptable JavaScript framework for building UI on the web.
-
Typescript
TypeScript is a superset of JavaScript that compiles to clean JavaScript output.
-
TensorFlow
An Open Source Machine Learning Framework for Everyone
-
Django
The Web framework for perfectionists with deadlines.
-
Laravel
A PHP framework for web artisans
-
D3
Bring data to life with SVG, Canvas and HTML. 📊📈🎉
-
Recommend Topics
-
javascript
JavaScript (JS) is a lightweight interpreted programming language with first-class functions.
-
web
Some thing interesting about web. New door for the world.
-
server
A server is a program made to process requests and deliver data to clients.
-
Machine learning
Machine learning is a way of modeling and interpreting data that allows a piece of software to respond intelligently.
-
Visualization
Some thing interesting about visualization, use data art
-
Game
Some thing interesting about game, make everyone happy.
Recommend Org
-
Facebook
We are working to build community through open source technology. NB: members must have two-factor auth.
-
Microsoft
Open source projects and samples from Microsoft.
-
Google
Google ❤️ Open Source for everyone.
-
Alibaba
Alibaba Open Source for everyone
-
D3
Data-Driven Documents codes.
-
Tencent
China tencent open source team.
from litstudy.