EgoTaskQA baselines
October 10, 2022 · View on GitHub
Data Download
For data and features download, please refer to our website.
Within the google drive, you can find features under (gdrive) features/ and question-answer pairs in (gdrive) data/qa.
You may as well generate all questions with codes in this repo. Here we clarify on the paths used for data storage:
After download, create $FEATURE_BASE under the current directory and put features, data and
checkpoitns to subdirectories as follows:
- download and put
(gdrive) features/video_feature_20.h5to$FEATURE_BASE. - download and put
(gdrive) features/lemma-qa_appearance_feat.h5,(gdrive) features/lemma-qa_motion_feat.h5(G-drive) to$FEATURE_BASE/hcrn_data/. - download
(gdrive) features/video_features.zipand unzip it to$FEATURE_BASE/video_features. - download
(gdrive) features/glove.840.300d.pklto$GLOVE_PT_PATH. - download and put
(gdrive) data/train_qas.json,(gdrive) data/test_qas.json,(gdrive) data/val_qas.json,(gdrive) data/tagged_qa.json,(gdrive) data/vid_intervals.jsonto$SAVE_BASE/$SPLIT.$SPLITcan bedirectorindirect. We will also discuss how to generate them from your own generated question-answer pairs as well in the following section.
Preprocessing
For running experiments, we need to first take all question-answer pairs generated and split them to train, validation and test sets. For this step, you can do the following:
$ cd baselines
$ python preprocess/qas2tagged_qas.py --file balance_qa_maskX.json --path $SAVE_BASE
$ python preprocess/split.py --type $SPLIT
where balance_qa_maskX.json and $SAVE_BASE are corresponding settings used in question-answer generation process.
$SPLIT can be direct or indirect depending on the split chosen. After these steps there should be (
taking direct as an example) the following in your $SAVE_BASE:
direct/
|-- tagged_qas.json
|-- test_qas.json
|-- train_qas.json
|-- val_qas.json
|-- vid_intervals.json
Next, After downloading all data to their correct locations, run the following for preprocessing:
$ chmod a+x preprocess.sh
$ ./preprocess.sh $SAVE_BASE $GLOVE_PT_PATH
After running this script, you should have the following in your direct directory:
direct/
|-- answer_set.txt
|-- all_reasoning_types.txt
|-- char_vocab.txt
|-- formatted_test_qas_encode.json
|-- formatted_train_qas_encode.json
|-- formatted_val_qas_encode.json
|-- glove.pt
|-- lemma-qa_vocab.json
|-- tagged_qas.json
|-- test_qas.json
|-- test_qas_encode.json
|-- train_qas.json
|-- train_qas_encode.json
|-- val_qas.json
|-- val_qas_encoder.json
|-- vid_intervals.json
|-- vocab.txt
which compared to the previous version, generates metadata files for experiments.
Training
Use the following command to train the model you want to experiment with (Specify $OUTPUT for logs):
# HCRN experiment
$ python train_hcrn.py --base_data_dir $SAVE_BASE/$SPLIT --feature_base $FEATURE_BASE/hcrn_data --basedir $OUTPUT
# HME or HGA ($TRAIN_MODEL_PY: train_hme.py, train_hga.py) experiment
$ python $TRAIN_MODEL_PY --base_data_dir $SAVE_BASE/$SPLIT --video_feature_path $FEATURE_BASE/video_feature_20.h5 --basedir $OUTPUT
# PSAC, LSTM, BERT, VisualBERT ($TRAIN_MODEL_PY: train_psac.py, train_pure_lstm.py, train_linguistic_bert.py, train_visual_bert.py) experiment
$ python $TRAIN_MODEL_PY --feature_base_path $FEATURE_BASE/video_features --base_data_dir $BASE_DATA_DIR --basedir $OUTPUT
For bert-based model, you need to set BertTokenizer_CKPT and BertModel_CKPT for the model to load pretrained model from huggingface.
- For linguistic_bert, set BertTokenizer_CKPT="bert-base-uncased", BertModel_CKPT="bert-base-uncased".
- For visual_bert, set BertTokenizer_CKPT="bert-base-uncased", VisualBertModel_CKPT="uclanlp/visualbert-vqa-coco-pre".
*Note, for ClipBERT related experiments, please follow the instruction from their website for preprocessing and training.
Reload ckpts & test_only
To reload checkpoints and only run inference on test_qas, run the following command:
# HCRN experiment
$ python train_hcrn.py --base_data_dir $SAVE_BASE/$SPLIT --feature_base $FEATURE_BASE/hcrn_data --reload_model_path $RELOAD_MODEL_PATH --test_only 1 --basedir $OUTPUT
# HME or HGA ($TRAIN_MODEL_PY:train_hme.py, train_hga.py) experiment
$ python $TRAIN_MODEL_PY --base_data_dir $SAVE_BASE/$SPLIT --video_feature_path $FEATURE_BASE/video_feature_20.h5 --reload_model_path $RELOAD_MODEL_PATH --test_only 1 --basedir $OUTPUT
# PSAC, LSTM, BERT, VisualBERT ($TRAIN_MODEL_PY: train_psac.py, train_pure_lstm.py, train_linguistic_bert.py, train_visual_bert.py) experiment
$ python $TRAIN_MODEL_PY --feature_base_path $FEATURE_BASE/video_features --base_data_dir $BASE_DATA_DIR --reload_model_path $RELOAD_MODEL_PATH --test_only 1 --basedir $OUTPUT
Acknowledgement
This code heavily used resources from VisualBERT, HCRN, HGA, HME, PSAC. We thank the authors for open-sourcing their awesome projects.