https://github.com/andreasmadsen/python-textualheatmap

Create interactive textual heat maps for Jupiter notebooks

https://github.com/andreasmadsen/python-textualheatmap

Science Score: 26.0%

This score indicates how likely this project is to be science-related based on various indicators:

  • CITATION.cff file
  • codemeta.json file
    Found codemeta.json file
  • .zenodo.json file
  • DOI references
    Found 1 DOI reference(s) in README
  • Academic publication links
  • Committers with academic emails
  • Institutional organization owner
  • JOSS paper metadata
  • Scientific vocabulary similarity
    Low similarity (9.8%) to scientific vocabulary
Last synced: 11 months ago · JSON representation

Repository

Create interactive textual heat maps for Jupiter notebooks

Basic Info
  • Host: GitHub
  • Owner: AndreasMadsen
  • License: mit
  • Language: Jupyter Notebook
  • Default Branch: master
  • Size: 6.24 MB
Statistics
  • Stars: 196
  • Watchers: 6
  • Forks: 14
  • Open Issues: 3
  • Releases: 0
Created over 6 years ago · Last pushed about 2 years ago
Metadata Files
Readme License

README.md

textualheatmap

Create interactive textual heatmaps for Jupiter notebooks.

I originally published this visualization method in my distill paper https://distill.pub/2019/memorization-in-rnns/. In this context, it is used as a saliency map for showing which parts of a sentence are used to predict the next word. However, the visualization method is more general-purpose than that and can be used for any kind of textual heatmap purposes.

textualheatmap works with python 3.6 or newer and is distributed under the MIT license.

Gif of saliency in RNN models

An end-to-end example of how to use the HuggingFace 🤗 Transformers python module to create a textual saliency map for how each masked token is predicted.

Open In Colab

Gif of saliency in BERT models

Install

bash pip install -U textualheatmap

API

Examples

Example of sequential-charecter model with metadata visible

Open In Colab

```python from textualheatmap import TextualHeatmap

data = [[ # GRU data {"token":" ", "meta":["the","one","of"], "heat":[1,0,0,0,0,0,0,0,0]}, {"token":"c", "meta":["can","called","century"], "heat":[1,0.22,0,0,0,0,0,0,0]}, {"token":"o", "meta":["country","could","company"], "heat":[0.57,0.059,1,0,0,0,0,0,0]}, {"token":"n", "meta":["control","considered","construction"], "heat":[1,0.20,0.11,0.84,0,0,0,0,0]}, {"token":"t", "meta":["control","continued","continental"], "heat":[0.27,0.17,0.052,0.44,1,0,0,0,0]}, {"token":"e", "meta":["context","content","contested"], "heat":[0.17,0.039,0.034,0.22,1,0.53,0,0,0]}, {"token":"x", "meta":["context","contexts","contemporary"], "heat":[0.17,0.0044,0.021,0.17,1,0.90,0.48,0,0]}, {"token":"t", "meta":["context","contexts","contentious"], "heat":[0.14,0.011,0.034,0.14,0.68,1,0.80,0.86,0]}, {"token":" ", "meta":["of","and","the"], "heat":[0.014,0.0063,0.0044,0.011,0.034,0.10,0.32,0.28,1]}, # ... ],[ # LSTM data # ... ]]

heatmap = TextualHeatmap( width = 600, showmeta = True, facettitles = ['GRU', 'LSTM'] )

Set data and render plot, this can be called again to replace

the data.

heatmap.set_data(data)

Focus on the token with the given index. Especially useful when

interactive=False is used in TextualHeatmap.

heatmap.highlight(159) ```

Shows saliency with predicted words at metadata

Example of sequential-charecter model without metadata

Open In Colab

When show_meta is not True, the meta part of the data object has no effect.

python heatmap = TextualHeatmap( facet_titles = ['LSTM', 'GRU'], rotate_facet_titles = True ) heatmap.set_data(data) heatmap.highlight(159)

Shows saliency without metadata

Example of non-sequential-word model

Open In Colab

format = True can be set in the data object to inducate tokens that are not directly used by the model. This is useful if word or sub-word tokenization is used.

```python data = [[ {'token': '[CLR]', 'meta': ['', '', ''], 'heat': [1, 0, 0, 0, 0, ...]}, {'token': ' ', 'format': True}, {'token': 'context', 'meta': ['today', 'and', 'thus'], 'heat': [0.13, 0.40, 0.23, 1.0, 0.56, ...]}, {'token': ' ', 'format': True}, {'token': 'the', 'meta': ['##ual', 'the', '##ually'], 'heat': [0.11, 1.0, 0.34, 0.58, 0.59, ...]}, {'token': ' ', 'format': True}, {'token': 'formal', 'meta': ['formal', 'academic', 'systematic'], 'heat': [0.13, 0.74, 0.26, 0.35, 1.0, ...]}, {'token': ' ', 'format': True}, {'token': 'study', 'meta': ['##ization', 'study', '##ity'], 'heat': [0.09, 0.27, 0.19, 1.0, 0.26, ...]} ]]

heatmap = TextualHeatmap(facettitles = ['BERT'], showmeta=True) heatmap.set_data(data) ```

Shows saliency in a BERT model, using sub-word tokenization

Citation

If you use this in a publication, please cite my Distill publication where I first demonstrated this visualization method.

bib @article{madsen2019visualizing, author = {Madsen, Andreas}, title = {Visualizing memorization in RNNs}, journal = {Distill}, year = {2019}, note = {https://distill.pub/2019/memorization-in-rnns}, doi = {10.23915/distill.00016} }

Sponsor

Sponsored by NearForm Research.

Owner

  • Name: Andreas Madsen
  • Login: AndreasMadsen
  • Kind: user
  • Location: Copenhagen, Denmark
  • Company: MILA

Researching interpretability for Machine Learning because society needs it.

GitHub Events

Total
Last Year

Committers

Last synced: over 1 year ago

All Time
  • Total Commits: 20
  • Total Committers: 1
  • Avg Commits per committer: 20.0
  • Development Distribution Score (DDS): 0.0
Past Year
  • Commits: 4
  • Committers: 1
  • Avg Commits per committer: 4.0
  • Development Distribution Score (DDS): 0.0
Top Committers
Name Email Commits
Andreas Madsen a****k@g****m 20

Issues and Pull Requests

Last synced: 11 months ago

All Time
  • Total issues: 6
  • Total pull requests: 2
  • Average time to close issues: about 1 year
  • Average time to close pull requests: 2 months
  • Total issue authors: 5
  • Total pull request authors: 2
  • Average comments per issue: 2.0
  • Average comments per pull request: 4.5
  • Merged pull requests: 1
  • Bot issues: 0
  • Bot pull requests: 0
Past Year
  • Issues: 0
  • Pull requests: 0
  • Average time to close issues: N/A
  • Average time to close pull requests: N/A
  • Issue authors: 0
  • Pull request authors: 0
  • Average comments per issue: 0
  • Average comments per pull request: 0
  • Merged pull requests: 0
  • Bot issues: 0
  • Bot pull requests: 0
Top Authors
Issue Authors
  • BigSalmon2 (2)
  • ScottishFold007 (1)
  • dr-smgad (1)
  • sadransh (1)
  • h-amirkhani (1)
Pull Request Authors
  • VenkteshV (1)
  • Willian-Zhang (1)
Top Labels
Issue Labels
Pull Request Labels

Dependencies

setup.py pypi
  • ipython *
package.json npm