AiRyunn/Pangu

Code for reproducing the ACL'23 paper: Don't Generate, Discriminate: A Proposal for Grounding Language Models to Real-World Environments

โ˜… 0Forks 0PythonGitHub โ†—Compare

Project website โ†—

README

๐Ÿ“ฃ ๐Ÿ“ฃ(This is mainly for reproducing the results in our ACL'2023 paper. We will release the generic Pangu library in OSU-NLP-Group/Pangu.)

Don't Generate, Discriminate: A Proposal for Grounding Language Models to Real-World Environments

Contributions Welcome License language-python3 made-with-Pytorch paper award

A key missing capacity of current language models (LMs) is grounding to real-world environments. Most existing work for grounded language understanding uses LMs to directly generate plans that can be executed in the environment to achieve the desired effects. It thereby casts the burden of ensuring grammaticality, faithfulness, and controllability all on the LMs. We propose Pangu, a generic framework for grounded language understanding that capitalizes on the discriminative ability of LMs instead of their generative ability. Pangu consists of a symbolic agent and a neural LM working in a concerted fashion: The agent explores the environment to incrementally construct valid plans, and the LM evaluates the plausibility of the candidate plans to guide the search process. A case study on the challenging problem of knowledge base question answering (KBQA), which features a massive environment, demonstrates the remarkable effectiveness and flexibility of Pangu: A BERT-base LM is sufficient for setting a new record on standard KBQA datasets, and larger LMs further bring substantial gains. Pangu also enables, for the first time, effective few-shot in-context learning for KBQA with large LMs such as Codex.

image image

Walk Through Pangu with KBQA

We instantiate Pangu on knowledge base question answering (KBQA), which is representative testbed for grounded language understanding with a highly complex and heterogeneous environment. pangu_compressed

File Structure

pangu/
โ”œโ”€  acl_configs/: configuration files for training and inference
โ”œโ”€  data/: KBQA data files (e.g., GrailQA)
โ”œโ”€  ontology/: Processed Freebase ontology files
โ”œโ”€  answer_typing/: Answer typing results
โ”œโ”€  el_results/: Entity linking results 
โ”œโ”€  utils/:
โ”‚    โ”œโ”€  bert_interface.py: Interface to BERT 
โ”‚    โ”œโ”€  huggingface_interface.py: Interface to Huggingface models 
โ”‚    โ”œโ”€  logic_form_util.py: Tools related to logical forms, including the exact match checker for two logical forms
โ”‚    โ”œโ”€  sparql_executor.py: Sparql-related tools
โ”‚    โ”œโ”€  kb_environment.py: Core functions for KB querying and constrained decoding
โ”‚    โ””โ”€โ”€ sparql_cache.py: Cache executions of different types of Sparql queries
โ”œโ”€  new_model/:
โ”‚    โ”œโ”€  bottom_up_parser.py: Pangu model class
โ”‚    โ””โ”€โ”€ bottom_up_parser_reader.py: Pangu dataset reader class
โ”œโ”€  new_model/: prediction results in json
โ”œโ”€  run.py: Main function
โ”œโ”€  trained_models.md: the links to our trained models; download them to make predictions
โ””โ”€โ”€ environment.yml: yml file for conda environment 

Results

Overall Results

image

Sample Efficiency

image

Strong Generalizability

image

Reproducing Our Results

Environment Setup

Please configure your own conda environment using environment.yml. Replace [your_conda_path] in that file with the path of your local anaconda folder, and then create the environment with conda env create -f environment.yml.

Training & Inference

PYTHONHASHSEED=23 python run.py \
    train \
    acl_configs/grail_train.jsonnet \
    --include-package \
    new_model.bottom_up_parser \
    --include-package \
    new_model.bottom_up_parser_reader \
    --include-package \
    utils.huggingface_interface \
    -s \
    [output_dir]

To train the model with multiple cards using DDP, uncomment the distributed field in the config file. Note that, training can be quite slow at an earlier stage, but it will be faster when more SPARQL queries are executed and cached.

To do inference with a saved model, use the first configuration in launch.json, or do

PYTHONHASHSEED=23 python run.py \
    predict \
    [output_dir]/model.tar.gz \
    [path_to_file] \
    --include-package \
    new_model.bottom_up_parser \
    --include-package \
    new_model.bottom_up_parser_reader \
    --include-package \
    utils.huggingface_interface \
    -output-file \
    predictions.txt \
    --use-dataset-reader \
    -o \
    "{'model': {'infer': true}, 'validation_dataset_reader': {'infer': true, 'perfect_entity_linking': false}}"

In utils.sparql_executer.py, replace "http://127.0.0.1:3094/sparql" with your own SPARQL endpoint.

Configure your experiments following configuration files under acl_configs.

Experiments with LLMs

Our original experiments with LLMs were done with Codex. However, since March 2023, Codex has been deprecated by OpenAI. We will adjust and upload this part of code soon.

Acknowledgements

The authors would like to thank Percy Liang, Jiawei Han, Jonathan Berant, Huan Sun, and other colleagues from the OSU NLP group for their valu- able feedback. The authors would also like to thank Shijie Chen and Chan Hee Song for proof-of-concept implementation of Pangu on other tasks, Yiheng Shu for sharing their entity linking results, and Tianbao Xie for clarifications on UnifiedSKG. This research was supported in part by ARL W911NF2220144, NSF OAC 2112606, and Ohio Supercomputer Center.

Citation

If you find our work helpful, please cite our ACL paper as follows.

@inproceedings{gu-etal-2023-dont,
    title = "Don{'}t Generate, Discriminate: A Proposal for Grounding Language Models to Real-World Environments",
    author = "Gu, Yu  and
      Deng, Xiang  and
      Su, Yu",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.270",
    pages = "4928--4949",
    abstract = "A key missing capacity of current language models (LMs) is grounding to real-world environments. Most existing work for grounded language understanding uses LMs to directly generate plans that can be executed in the environment to achieve the desired effects. It thereby casts the burden of ensuring grammaticality, faithfulness, and controllability all on the LMs. We propose Pangu, a generic framework for grounded language understanding that capitalizes on the discriminative ability of LMs instead of their generative ability. Pangu consists of a symbolic agent and a neural LM working in a concerted fashion: The agent explores the environment to incrementally construct valid plans, and the LM evaluates the plausibility of the candidate plans to guide the search process. A case study on the challenging problem of knowledge base question answering (KBQA), which features a massive environment, demonstrates the remarkable effectiveness and flexibility of Pangu: A BERT-base LM is sufficient for setting a new record on standard KBQA datasets, and larger LMs further bring substantial gains.Pangu also enables, for the first time, effective few-shot in-context learning for KBQA with large LMs such as Codex.",
}

Contributors

entslscheiaysu1989

Issues