This repository holds code dedicated to extracting some knowledge graph our of the 1957 pdf version of Paul Reps' compilation Zen flesh, zen bones.
- Notes concerning the original book
- Running the converter
- Running the full data workflow with jejune_cli
This directory holds a copy of the 1957 pdf version of Paul Reps' compilation Zen flesh, zen bones as offered online by OceanoPDF.
- Within chapter
101 ZEN STORIES,- there are two sub-chapters with number
46. - there is no sub-chapter numbered
49.
- there are two sub-chapters with number
- Within chapter the
CENTERINGchapter, there is a numbered list entry labeled65. blah-blah without any a or mthat misses its ending dot character.. Alas such an entry matches the chapter pattern which fools the documentation reconstruction into believing this is a real chapter... - They are many typos in the original text e.g. look for occurrences of
yonin place ofyouin chapter2. Finding a Diamond on a Muddy Road... Refer totypo_and_fixentries withinConvert/StructuralInfo.py.
cd `git rev-parse --show-toplevel`/Convert
python3.10 -m venv venv
source ./venv/bin/activate
pip install -r requirements.txt
python main.pyWithin the above running context (directory and installed virtual environment)
cd `git rev-parse --show-toplevel`/Convert
pip install -r requirements-dev.txt
pytest test_main.pyOnce development has improved some resulting converted files the following command will overwrite the reference resulting data
python main.py --output_directory ../result_data/docker build -t jejune:doc_Zen_Flesh_Zen_Bones https://github.com/EricBoix/jj_doc_Zen_Flesh_Zen_Bones.git#:DockerContext
docker run --rm jejune:doc_Zen_Flesh_Zen_Bones --helpExtracting the result out of the container requires local filesystem mount
docker run --rm -v `pwd`/junk:/output jejune:doc_Zen_Flesh_Zen_Bones --output_directory /outputInstall and configure jejune_cli, then run jejune doctor to verify the configuration. This boils down to
uv tool install git+https://github.com/EricBoix/jejune_cli
jejune configuration init
# Proceed with the configuration of the files located in .jejune/ sub-directory.
# Assert the configuration is sound with
jejune doctorDefine a convenience variable for the results directory:
export RESULTS_DIR=`pwd`/result_dataRun the converter to extract a markdown out of the original PDF :
jejune convert build
jejune convert run --output-dir $RESULTS_DIRRun the (Knowledge Graph) extraction (starting a neo4j database being prerequisite)
jejune neo4j delete $RESULTS_DIR # Avoid collision with previous/other run
jejune neo4j stats --assert 0/0 # Just making sure deletion was effective
jejune neo4j start $RESULTS_DIR
jejune graph extract $RESULTS_DIR \
--load_markdown_document \
1957_-_Paul_Reps_-_Zen_flesh_zen_bones-A_Collection_of_Zen_and_Pre_Zen_Writings_-_Scan_by_OceanofPDF_dot_com_-_local_converter.md \
--load_json_document \
1957_-_Paul_Reps_-_Zen_flesh_zen_bones-A_Collection_of_Zen_and_Pre_Zen_Writings_-_Scan_by_OceanofPDF_dot_com_-_Sentences_as_LangChain_Document.json
jejune neo4j stats --assert 4769/9962Optional: dump the database content for later usage (and restore it to assert dump integrity/validity)
jejune neo4j stop
jejune neo4j dump $RESULTS_DIR neo4j.ZenFleshZenBones.MarkdownTextSplitterAndSentences.dump
# Restore the database out of the dump (just to make sure)
# WARNING: restoring DELETEs the existing database
jejune neo4j restore $RESULTS_DIR neo4j.ZenFleshZenBones.MarkdownTextSplitterAndSentences.dump
jejune neo4j start $RESULTS_DIR
jejune neo4j stats --assert 4769/9962Extract knowledge graph in Turtle format
jejune neo4j dump-turtle $RESULTS_DIR ZenFleshZenBones.MarkdownTextSplitterAndSentences.ttl
jejune neo4j stop