Niah (Statistical Disclosure Control) Wiki
Niah supports k-anonymity statistical disclosure control assessments
Brought to you by:
countwitte
Welcome to your wiki!
This is the default page, edit it as you see fit. To add a new page simply reference it within brackets, e.g.: [SamplePage].
The wiki uses Markdown syntax.
Documentation
** Needle in a Haystack (Niah) Algorithm **
Copyright (C) 2012 Michael Comerford
comments/questions: michael@secretplan4b.co.uk
This program was developed as part of PhD Research at the University of Glasgow, and work carried out on internship with the Scottish Government.
Niah takes data formatted as comma separated values (csv) along with, a set of key variables (given by column number, starting from 0), and a threshold value k. Niah splits the file into 'atrisk' and 'safe' files based on an an assessment of k-anonymity (all records should be indistinguishable from k-1 other records).
This tool should be used as a part of a comprehensive statistical disclosure risk assessment. This should include: consideration of the data environment that data or statistical results are being published into; an assessment of how sensitive the data is, and the impact of any disclosure. These two aspects will assist in assigning an appropriate threshold level k.
As an example of k as defined above, if -th is set to 5 all combinations of the key variables with a record count of 4 or less will be deemed at risk.
Due to the memory requirements for processing large csv files, you may need to increase the maximum memory used by the Java Virtual Machine (use the -Xmx[value][unit] option, e.g. -Xmx512m) A general rule is that Niah requires the csv file x 10 available memory.
USAGE: Niah -kv "VAR1,...,VARx" [-th N] [-o OUTPUTNAME] FILE
OPTIONS:
-kv "VAR1,...,VARx" specify the column number of the key variables (comma separated)
-th N (optional) specify the threshold for k-anonymity (defaults to 1)
-o OUTPUTNAME (optional) specify the output file names (without extension)
FILE specify the filename for the input data (csv format required)
STATA Example of Niah implementation
Example usage for survey data originally in Stata format:
*** Using Niah
** Data setup:
use asex aage ajbsoc aqfedhi aregion ///
using c:\data\bhps\bh11810\aindresp.dta, clear
describe /* 10264 cases */
outsheet using C:\soft\niah\data\bhps_w1_eg1.csv, comma nolabel replace
**
** Example of testing a k-anonymity at threshold 3 for age and gender in Niah
/*
MSDOS> cd soft\niah\Niah
MSDOS> java -version #java required should be 1.6
MSDOS> java -Xmx1g Niah -kv "0,3" -th 3 -o results ..\data\bhps_w1_eg1.csv
The key variable indices are: [0, 3]
The threshold for k-anonymity test is set at: 3
Data in = bhps_w1_eg1.csv
safe data = results_safe.csv
at risk data = results_atrisk.csv
*/
insheet using C:\soft\niah\Niah\results_safe.csv, clear
summarize /* 10256 cases 'safe' under the threshold specified */
insheet using C:\soft\niah\Niah\results_atrisk.csv, clear
summarize /* 8 cases 'safe' under the threshold specified */