naaccr

Read cancer records in the NAACCR format

https://github.com/werthpadoh/naaccr

Science Score: 26.0%

This score indicates how likely this project is to be science-related based on various indicators:

  • CITATION.cff file
  • codemeta.json file
    Found codemeta.json file
  • .zenodo.json file
    Found .zenodo.json file
  • DOI references
  • Academic publication links
  • Committers with academic emails
  • Institutional organization owner
  • JOSS paper metadata
  • Scientific vocabulary similarity
    Low similarity (13.7%) to scientific vocabulary

Keywords

naaccr rstats
Last synced: 11 months ago · JSON representation

Repository

Read cancer records in the NAACCR format

Basic Info
  • Host: GitHub
  • Owner: WerthPADOH
  • License: other
  • Language: R
  • Default Branch: master
  • Size: 1.21 MB
Statistics
  • Stars: 10
  • Watchers: 2
  • Forks: 2
  • Open Issues: 3
  • Releases: 0
Topics
naaccr rstats
Created about 8 years ago · Last pushed about 1 year ago
Metadata Files
Readme Changelog License

README.Rmd

---
title: "naaccr"
output:
  github_document:
    html_preview: false
---

```{r setup, include=FALSE}
knitr::opts_chunk$set(
  collapse = TRUE,
  comment  = "#>",
  fig.path = "man/figures/README-"
)
```

## Summary

The `naaccr` R package enables researchers to easily read and begin analyzing
cancer incidence records stored in the
[North American Association of Central Cancer Registries](https://www.naaccr.org/)
(NAACCR) file format.

## Usage

`naaccr` focuses on two tasks: arranging the records and preparing the fields
for analysis.

### Records

The `naaccr_record` class defines objects which store cancer incidence records.
It inherits from `data.frame`, and for now only makes sure a dataset has a
standard set of columns. While `naaccr_record` has a singular-sounding name, it
can contain multiple records as rows.

The `read_naaccr` function creates a `naaccr_record` object from a
NAACCR-formatted file.

```{r showRecords}
record_file <- system.file(
  "extdata/synthetic-naaccr-18-abstract.txt",
  package = "naaccr"
)
record_lines <- readLines(record_file)
## Marital status and race fields
cat(substr(record_lines[1:5], 206, 216), sep = "\n")
```

```{r readNaaccr}
library(naaccr)

records <- read_naaccr(record_file, version = 18)
records[1:5, c("maritalStatusAtDx", "race1", "race2", "race3")]
```

By default, `read_naaccr` reads all fields defined in a format. For example,
the NAACCR 18 format used above has `r nrow(naaccr_format_18)` fields. Rarely
would an analysis need even 100 fields. By specifying which fields to keep, one
can improve time and memory efficiency.

```{r readKeepColumns}
dim(records)
records_slim <- read_naaccr(
  input       = record_file,
  version     = 18,
  keep_fields = c("ageAtDiagnosis", "countyAtDx", "primarySite")
)
dim(records_slim)
```

Like with most classes, one can create a new `naaccr_record` object with the
function of the same name. The result will have the given columns.

```{r naaccrRecord}
nr <- naaccr_record(
  primarySite = "C010",
  dateOfBirth = "19450521"
)
nr[, c("primarySite", "dateOfBirth")]
```

The `as.naaccr_record` function can transform an existing data frame. It does
require any existing columns to use NAACCR's XML names.

```{r asNaaccrRecord}
prefab <- data.frame(
  ageAtDiagnosis = c(1, 120, 999),
  race1          = c("01", "02", "88")
)
converted <- as.naaccr_record(prefab)
converted[, c("ageAtDiagnosis", "race1")]
```

### Code translation

The NAACCR format uses similar schemes for a lot of fields, and the `naaccr`
package includes functions to help translate them.

`naaccr_boolean` translates "yes/no" fields. By default, it assumes `"0"` stands
for "no", and `"1"` stands for "yes."

```{r naaccrBoolean}
naaccr_boolean(c("0", "1", "2"))
```

Some fields use `"1"` for `FALSE` and `"2"` for `TRUE`. Use the `false_value`
parameter to work with these.

```{r falseValue}
naaccr_boolean(c("0", "1", "2"), false_value = "1")
```

#### Categorical fields

The `naaccr_factor` function translates values using a specific field's category
codes.

```{r naaccrFactor}
naaccr_factor(c("01", "31", "65"), "primaryPayerAtDx")
```

Some fields have multiple codes explaining why an actual value isn't known.
By default, they'll all be converted to `NA` so they can propagate that information in R.
But the reasons can be useful, so `naaccr_factor` and `naaccr_record` both have
a `keep_unknown` parameter.

```{r keepUnknown}
naaccr_factor(c("1", "9"), field = "sex")
naaccr_factor(c("1", "9"), field = "sex", keep_unknown = TRUE)
naaccr_record(
  sex = c("1", "9"),
  race1 = c("01", "99"),
  keep_unknown = TRUE,
  version = "25"
)
```

#### Numeric with special missing

Some fields contain primarily continuous or count data but also use special
codes. One name for this type of code is a "sentinel value." The
`split_sentineled` function splits these fields in two.

```{r naaccrSentineled}
rnp <- split_sentineled(c(10, 20, 90, 95, 99, NA), "regionalNodesPositive")
rnp
```

## Building

```{r needForBuild}
library(devtools)

deps <- packageDescription("naaccr", fields = c("Depends", "Imports", "Suggests"))
deps <- Filter(function(x) any(!is.na(x)), deps)
dep_names <- lapply(deps, function(x) devtools::parse_deps(x)[["name"]])
dep_names <- sort(unlist(dep_names))
dep_list <- paste0("- `", dep_names, "`", collapse = "\n")
```

To build the `naaccr` package, you'll need the following R packages:

`r dep_list`

To document, build, and test the package, run the `build.R` script with the
package's root as the working directory.

## Project files

First, know this project fills two roles:

1.  Creating a package to work with NAACCR data in R.
2.  Collecting the data needed to process NAACCR files in plain-text and
    machine-readable formats.

```
naaccr/
├ R/                  # R files to create the package objects
├ data-raw/           # Plain-text data files and scripts for processing them
│ ├ code-labels/      # Mappings of codes to understandable labels
│ ├ sentinel-labels/  # Mappings of sentinel values to understandable labels
│ └ record-formats/   # Tables defining each NAACCR file format
├ external/           # Downloaded files and scripts to create files in `data-raw`
├ inst/
│ └ extdata/          # Data files for examples in the documentation
└ tests/              # tests and data using the `testthat` package
```

Files in `external` only need to be updated or run when NAACCR publishes a new
or revised format. In that case, refer to the comments in the `.R` scripts in
that directory for where to download the new files.

Think of these scripts as handy tools for generating `data-raw` files.
Some cleaning of their output may be required.

To run `create-record-format-files.R`, you'll need to create an account for the
[SEER API](https://api.seer.cancer.gov/) from the National Cancer Institute's
Surveillance, Epidemiology and End Results (SEER) program.
Store the API key as an environment variable named `SEER_API_KEY`.

Owner

  • Name: Nathan Werth
  • Login: WerthPADOH
  • Kind: user
  • Location: York, Pennsylvania
  • Company: Pennsylvania Department of Health

GitHub Events

Total
Last Year

Committers

Last synced: over 3 years ago

All Time
  • Total Commits: 411
  • Total Committers: 1
  • Avg Commits per committer: 411.0
  • Development Distribution Score (DDS): 0.0
Top Committers
Name Email Commits
Nathan Werth n****h@p****v 411
Committer Domains (Top 20 + Academic)
pa.gov: 1

Packages

  • Total packages: 1
  • Total downloads:
    • cran 296 last-month
  • Total dependent packages: 0
  • Total dependent repositories: 0
  • Total versions: 5
  • Total maintainers: 1
cran.r-project.org: naaccr

Read Cancer Records in the NAACCR Format

  • Versions: 5
  • Dependent Packages: 0
  • Dependent Repositories: 0
  • Downloads: 296 Last month
Rankings
Stargazers count: 16.3%
Forks count: 17.8%
Average: 28.7%
Dependent packages count: 29.8%
Dependent repos count: 35.5%
Downloads: 44.1%
Maintainers (1)
Last synced: 12 months ago

Dependencies

DESCRIPTION cran
  • data.table * imports
  • stringi * imports
  • utils * imports
  • ISOcodes * suggests
  • devtools * suggests
  • httr * suggests
  • jsonlite * suggests
  • magrittr * suggests
  • rmarkdown * suggests
  • roxygen2 * suggests
  • rvest * suggests
  • testthat * suggests
  • xml2 * suggests