https://github.com/britishgeologicalsurvey/tweet_classification
hazard-event classification from tweets
https://github.com/britishgeologicalsurvey/tweet_classification
Science Score: 10.0%
This score indicates how likely this project is to be science-related based on various indicators:
-
○CITATION.cff file
-
○codemeta.json file
-
○.zenodo.json file
-
○DOI references
-
✓Academic publication links
Links to: zenodo.org -
○Academic email domains
-
○Institutional organization owner
-
○JOSS paper metadata
-
○Scientific vocabulary similarity
Low similarity (8.8%) to scientific vocabulary
Keywords
Repository
hazard-event classification from tweets
Basic Info
Statistics
- Stars: 1
- Watchers: 1
- Forks: 1
- Open Issues: 2
- Releases: 0
Topics
Metadata Files
README.md
Automatic tweet classification for geohazards
This repository contains several techniques that can be used to perform binary classification with text. In this case we train each classifier to learn whether a tweet is referring to a hazard-event/phenomenon (earthquakes, volcanoes, floods, the aurora) or not.
The methods used in this implementation are the following:
- Convolutional Neural Network
- Recurrent Neural Network
- Baseline RNN
- GRU
- LSTM
- SVM
Prerequisites
- Python 3
Install dependencies
pip3 install requirements.txt
1. Neural Networks
1.1 Training
There are two modes of training for each algorithm:
- Training with random word vectors
- Training with pre-trained word2vec vectors
- Random-word vectors for CNN
python train_cnn.py
To train the RNN-LSTM with random word vectors, we have to use the following flag
python train_rnn.py --cell_type "LSTM"
using python train_rnn.py without a flag will train the baseline RNN.
- Pre-trained word vectors (word2vec):
python train_rnn.py --word2vec 'word2vec_twitter_model.bin'
After the training finishes the top five models with the best accuracy
are saved in a folder with the format runs/DAY_MMM_DD_HH_MM_SS_YYYY
1.2 Evaluation
For the evaluation we have to use the following flag to point to the directory that we saved the checkpoints
python eval.py --checkpoint_dir 'runs/DAY_MMM_DD_HH_MM_SS_YYYY/models'
This function will print the metrics that we use to measure the performance of the classifier.
## Accuracy: 0.863636
## F1: 0.8683602771362586
## Specificity: 0.79
## Precision: 0.8744186046511628
## Recall: 0.8623853211009175
To use tensorboard
tensorboard --logdir runs/DAY_MMM_DD_HH_MM_SS_YYYY/summaries
2. SVM
2.1 Training
There is more than one choice for a classifier. In order to train the classifier that is selected
python classifier_svm.py
The output will be a confusion matrix of the classifier based on the training data.
2.2 Evaluation
In order to evaluate the selected algorithm
python apply_classifier.py
The result will be a confusion matrix of the classifier based on the test data along with the accuracy.
Owner
- Name: British Geological Survey
- Login: BritishGeologicalSurvey
- Kind: organization
- Email: enquiries@bgs.ac.uk
- Location: Keyworth, Nottinghamshire
- Website: https://www.bgs.ac.uk
- Twitter: BritGeoSurvey
- Repositories: 71
- Profile: https://github.com/BritishGeologicalSurvey
The British Geological Survey is responsible for advising the UK government on geoscience and providing impartial advice to industry, academia and the public.
GitHub Events
Total
Last Year
Dependencies
- numpy ==1.15.1
- scikit_learn ==0.20.1
- tensorflow ==1.15.4