This system is built on the original JAMR code. The original JAMR code has its readme copied below. READ THAT FIRST.
Assuming you've read the JAMR docs, here's how to run the Robust Subgraph Generation parser:
First follow the instructions in the JAMR section on downloading all of its dependencies, then use the appropriate scripts/config_*.sh to set up environment variables. Then run scripts/EVAL.sh script, which we've modified to do a comparison between our system and two variants of JAMR.
This will run 3 different systems for comparison, JAMR, JAMR + Stanford Subgraph Generation, and JAMR + Gold Subgraphs. The Stanford Subgraph output will be titled *.parsed-stanford-concepts. The system should then also output a *.results file, assuming everything was set up correctly, which will list smatch scores (using the JAMR version of the smatch script, which also reports precision and recall).
#Hacking
DISCLAIMER: This is academic code, it's messy.
The entry point for JAMR is AMRParser.scala. It's been modified to accept the flag "--stanford-chunk-gen". This will call into StanfordDecoder.decode() during the NER++ stage, which will in turn create an instance of nlp.experiments.SequenceSystem, which will load two models: a sequence tagger, and a dictionary tag chunk classifier. Pre-trained models (LDC2013E117), serialized, are included in data/deft-train-manygen-classifier.ser.gz and data/deft-train-seq-classifier.ser.gz. If no other configuration is provided, the SequenceSystem will automatically load these models. SequenceSystem is a good place to start if you want to hack on the NER++ components. If you're interested in improving SRL++, start in AMRParser.scala.
#To Insert Your Own NER++ System
To implement your own NER++ system to compare against ours, we recommend taking the following steps:
- Provide your own Decoder.scala implementation (look at StanfordDecoder/Decoder.scala for reference)
- Add a flag to the flag parser in AMRParser.scala, ours is on line 36
- Modify the if case on line 209 of AMRParser.scala to use your decoder if your decoder flag (from the last step) is present
- Modify scripts/EVAL.sh to run another test using your system. Our test call is on lines 75-89, you can copy paste that and modify the call by replacing "--stanford-chunk-gen" with your own flag, and changing the "*.parsed-stanford-concepts" and "*.parsed-stanford-concepts.err" to your own unique extension
- Copy lines 115-118 of scripts/EVAL.sh to run smatch against the output files you generated in the code your wrote in the last step, and tee it into the *.results file
#To Insert Your Own SRL++ System
Good luck! That's all JAMR code, so crack out your Scala chops and figure it out. We still don't totally understand how it works, so won't be able to help you there.
JAMR is a semantic parser and aligner for the Abstract Meaning Representation.
We have released hand-alignments for 200 sentences of the AMR corpus.
For the performance of the parser, see docs/Parser_Performance.
#Building
First checkout the github repository (or download the latest release):
git clone https://github.com/jflanigan/jamr.git
JAMR depends on Scala, Illinois NER
system v2.7, tokenization scripts in
cdec, and WordNet for the
aligner. To download these dependencies into the subdirectory tools, cd to the jamr repository and run (requires
wget to be installed):
./setup
You should agree to the terms and conditions of the software dependencies before running this script. If you download
them yourself, you will need to change the relevant environment variables in scripts/config.sh. You may need to edit
the Java memory options in the script sbt and build.sbt if you get out of memory errors.
Source the config script - you will need to do this before running any of the scripts below:
. scripts/config.sh
Run ./compile to build an uberjar, which will be output to
target/scala-{scala_version}/jamr-assembly-{jamr_version}.jar (the setup script does this for you).
#Running the Parser
Download and extract model weights models.tgz into the directory
$JAMR_HOME/models. To parse a file (cased, untokenized, with one sentence per line) with the model trained on
LDC2014E41 data do:
. scripts/config.sh
scripts/PARSE.sh < input_file > output_file 2> output_file.err
The output is AMR format, with some extra fields described in docs/Nodes and Edges
Format and docs/Alignment Format. To run the parser trained
on other datasets (such as the older LDC2013E117 data, or freely downloadable Little
Prince data) source the config scripts config_LDC203E41.sh
or config_Little_Prince.sh instead.
#Running the Aligner
To run the rule-based aligner:
. scripts/config.sh
scripts/ALIGN.sh < amr_input_file > output_file
The output of the aligner is described in docs/Alignment Format. Currently the aligner works best for release r3 data (AMR Specification v1.0), but it will run on newer data as well.
#Hand Alignments
To create the hand alignments file, see docs/Hand Alignments.
#Experimental Pipeline
The following describes how to train and evaluate the parser. There are scripts to train the parser on various datasets, as well as a general train script to train the parser on any AMR dataset. More detailed instructions for training the parser are in docs/Step by Step Training.
To train the parser on LDC data or public AMR Bank data, download the data .tgz file
into to $JAMR_HOME/data/ and run one of the train scripts. The data file and the train script to run for each of the datasets
is listed in the following table:
| Dataset | Date released | Size (# sents) | Script to run | File to move to data/ |
|---|---|---|---|---|
| LDC2014T12 | June 16, 2014 | 13,051 | scripts/train_LDC2014T12.sh |
amr_anno_1.0_LDC2014T12.tgz |
| LDC2014E41 | May 30, 2014 | 18,779 | scripts/train_LDC2014E41.sh |
LDC2014E41_DEFT_Phase_1_AMR_Annotation_R4.tgz |
| LDC2013E117 (Proxy only) | October 14, 2013 | 8,219 | scripts/train_LDC2013E117.sh |
LDC2013E117.tgz |
| AMR Bank v1.4 | November 14, 2014 | 1,562 | scripts/train_Little_Prince.sh |
(automatically downloaded) |
For LDC2013E117 or LDC2014E41, you will need a license for LDC DEFT project data. The trained model will go into a subdirectory of models/ and the evaulation results will be printed and saved to
models/directory/RESULTS.txt. The performance of the parser on the various datasets is in docs/Parser
Performance.
To train the parser on another dataset, create a config file in scripts/ and
then do:
. scripts/my_config_file.sh
scripts/TRAIN.sh
The trained model will be saved into the $MODEL_DIR specified in the config script, and the results saved in
$MODEL_DIR/RESULTS.txt To run the parser with your trained model, source my_config_file.sh before running
PARSE.sh.
To evaluate a trained model against a gold standard AMR file, do:
. scripts/my_config_file.sh
scripts/EVAL.sh gold_amr_file
The predicted output will be in models/my_directory/gold_amr_file.parsed-gold-concepts for the parser with oracle
concept ID, models/my_directory/gold_amr_file.parsed for the full pipeline, and the results saved in
models/my_directory/gold_amr_file.results.