NLP_example.md
January 8, 2025 ยท View on GitHub
NLP Example: Vectorization and Visualization of Math Papers on arXiv
In this example, we analyze math papers on arXiv to demonstrate Natural Language Processing (NLP) techniques.
Note: This example runs only locally and is not compatible with Google Colaboratory.
While the parameters have not been fine-tuned, this serves as a starting point for more serious NLP applications.
Steps to Run the Example
-
Install Required Libraries
Ensure you havegensim,nltk, andkmapperinstalled:conda install gensim nltk pip install kmapper -
Install metha
For Linux users: Download the binary release and install it:
sudo apt install ./metha_?.?.??_amd64.deb
- Download Metadata from arXiv
Use the following command to synchronize metadata from arXiv (this may take time):
metha-sync -format arXiv -set math -base-dir .metha http://export.arxiv.org/oai2
- Learn Vector Representations and Visualize
Generate a vector representation of papers using Doc2Vec and visualize it with Mapper:
python metha2df.py -o arxiv_mapper.html
The first run creates a pickled DataFrame (math_2007.pkl) containing metadata, along with Doc2Vec model files:
-- doc2vec_arxiv.model
-- doc2vec_arxiv.model.docvecs.vectors_docs.npy
This process may take a long time.
From the second time, the pre-trained model will be loaded, and only the visualization will be generated.
- Open the Visualization
Open arxiv_mapper.html in your browser to explore the results.