๐ฃ ๐ฃ(This is mainly for reproducing the results in our ACL'2023 paper. We will release the generic Pangu library in OSU-NLP-Group/Pangu.)
A key missing capacity of current language models (LMs) is grounding to real-world environments. Most existing work for grounded language understanding uses LMs to directly generate plans that can be executed in the environment to achieve the desired effects. It thereby casts the burden of ensuring grammaticality, faithfulness, and controllability all on the LMs. We propose Pangu, a generic framework for grounded language understanding that capitalizes on the discriminative ability of LMs instead of their generative ability. Pangu consists of a symbolic agent and a neural LM working in a concerted fashion: The agent explores the environment to incrementally construct valid plans, and the LM evaluates the plausibility of the candidate plans to guide the search process. A case study on the challenging problem of knowledge base question answering (KBQA), which features a massive environment, demonstrates the remarkable effectiveness and flexibility of Pangu: A BERT-base LM is sufficient for setting a new record on standard KBQA datasets, and larger LMs further bring substantial gains. Pangu also enables, for the first time, effective few-shot in-context learning for KBQA with large LMs such as Codex.
We instantiate Pangu on knowledge base question answering (KBQA), which is representative testbed for grounded language understanding with a highly complex and heterogeneous environment.

pangu/
โโ acl_configs/: configuration files for training and inference
โโ data/: KBQA data files (e.g., GrailQA)
โโ ontology/: Processed Freebase ontology files
โโ answer_typing/: Answer typing results
โโ el_results/: Entity linking results
โโ utils/:
โ โโ bert_interface.py: Interface to BERT
โ โโ huggingface_interface.py: Interface to Huggingface models
โ โโ logic_form_util.py: Tools related to logical forms, including the exact match checker for two logical forms
โ โโ sparql_executor.py: Sparql-related tools
โ โโ kb_environment.py: Core functions for KB querying and constrained decoding
โ โโโ sparql_cache.py: Cache executions of different types of Sparql queries
โโ new_model/:
โ โโ bottom_up_parser.py: Pangu model class
โ โโโ bottom_up_parser_reader.py: Pangu dataset reader class
โโ new_model/: prediction results in json
โโ run.py: Main function
โโ trained_models.md: the links to our trained models; download them to make predictions
โโโ environment.yml: yml file for conda environment
Please configure your own conda environment using environment.yml. Replace [your_conda_path] in that file with the path of your local anaconda folder, and then create the environment with conda env create -f environment.yml.
PYTHONHASHSEED=23 python run.py \
train \
acl_configs/grail_train.jsonnet \
--include-package \
new_model.bottom_up_parser \
--include-package \
new_model.bottom_up_parser_reader \
--include-package \
utils.huggingface_interface \
-s \
[output_dir]
To train the model with multiple cards using DDP, uncomment the distributed field in the config file.
Note that, training can be quite slow at an earlier stage, but it will be faster when more SPARQL queries are executed and cached.
To do inference with a saved model, use the first configuration in launch.json, or do
PYTHONHASHSEED=23 python run.py \
predict \
[output_dir]/model.tar.gz \
[path_to_file] \
--include-package \
new_model.bottom_up_parser \
--include-package \
new_model.bottom_up_parser_reader \
--include-package \
utils.huggingface_interface \
-output-file \
predictions.txt \
--use-dataset-reader \
-o \
"{'model': {'infer': true}, 'validation_dataset_reader': {'infer': true, 'perfect_entity_linking': false}}"
In utils.sparql_executer.py, replace "http://127.0.0.1:3094/sparql" with your own SPARQL endpoint.
Configure your experiments following configuration files under acl_configs.
Our original experiments with LLMs were done with Codex. However, since March 2023, Codex has been deprecated by OpenAI. We will adjust and upload this part of code soon.
The authors would like to thank Percy Liang, Jiawei Han, Jonathan Berant, Huan Sun, and other colleagues from the OSU NLP group for their valu- able feedback. The authors would also like to thank Shijie Chen and Chan Hee Song for proof-of-concept implementation of Pangu on other tasks, Yiheng Shu for sharing their entity linking results, and Tianbao Xie for clarifications on UnifiedSKG. This research was supported in part by ARL W911NF2220144, NSF OAC 2112606, and Ohio Supercomputer Center.
If you find our work helpful, please cite our ACL paper as follows.
@inproceedings{gu-etal-2023-dont,
title = "Don{'}t Generate, Discriminate: A Proposal for Grounding Language Models to Real-World Environments",
author = "Gu, Yu and
Deng, Xiang and
Su, Yu",
booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
month = jul,
year = "2023",
address = "Toronto, Canada",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2023.acl-long.270",
pages = "4928--4949",
abstract = "A key missing capacity of current language models (LMs) is grounding to real-world environments. Most existing work for grounded language understanding uses LMs to directly generate plans that can be executed in the environment to achieve the desired effects. It thereby casts the burden of ensuring grammaticality, faithfulness, and controllability all on the LMs. We propose Pangu, a generic framework for grounded language understanding that capitalizes on the discriminative ability of LMs instead of their generative ability. Pangu consists of a symbolic agent and a neural LM working in a concerted fashion: The agent explores the environment to incrementally construct valid plans, and the LM evaluates the plausibility of the candidate plans to guide the search process. A case study on the challenging problem of knowledge base question answering (KBQA), which features a massive environment, demonstrates the remarkable effectiveness and flexibility of Pangu: A BERT-base LM is sufficient for setting a new record on standard KBQA datasets, and larger LMs further bring substantial gains.Pangu also enables, for the first time, effective few-shot in-context learning for KBQA with large LMs such as Codex.",
}