Version Notes
This explains changes in any new versions of the data. Note that the version numbers for the latest versions of NIDS have not been updated in the do files in the program library that has been made available for data manipulation. Data users should update the global to the version of the data they are using.
Version 4 of the NIDS Wave 1 data, produced 20 February 2012, had the following changes from version 3:
1. Sample size (one household dropped).
Household 109011 has been dropped from the sample. This household was interviewed twice under phase 1 (with hhid 109009 ) and under phase 2 ( with hhid 109011). The latter interview has been excluded from the sample. As a result there is 1 fewer household in the HouseholdQ file. The two individuals (an adult and a proxy) in this household have been dropped from the HouseholdRoster file and the individual level files.
2. Duplicate record for polygamous relationship in HouseholdRoster file
A duplicate record has been created in the HouseholdRoster file for pid 316891 with different hhid's. This individual is resident in two households (106506 and 106744). The implication of this is that individuals in the HouseholdRoster file are now uniquely identified by hhid and pid, rather than pid only. This will be useful to remember when it comes to merging the different STATA files together.
Note that because of the polygamist's pid appearing in two different households, merges should take place on both hhid and pid in Version 4.0 of the data.
3. New Variables added
The following variables have been added to the public release dataset:
a. Interviewer Evaluation variables have been included in the Adult, Child, Proxy and HouseholdQ files.
b. New pid variables have been added for every question that asks for the pcode of the respondent.
c. Date of interview day has been included in all datasets. Users have found this variable useful, hence its inclusion.
4. Variables dropped
a.The household questionnaire variable, hhqi, has been dropped from the wave1 dataset, to avoid confusion with the wave 2 hhid variable in terms of content. The hhqi variable also no longer appears in the wave 2 dataset.
b. Hhderived weight variables have been dropped
5. Variable names updated
a. Year of interview has been included in Education variable names in the Adult, Child and Proxy files
b. The hhid variable has been renamed to w1_hhid in all datasets
6. Update of Key variables, Weights and Z-scores
In the process of undertaking Wave 2 fieldwork, key variables (dob, age, gender, education and race) have been updated as better responses from respondents were received in cases where there had been non responses or errors in wave 1 data. As a result of changes in these key variables and samples sizes of the dataset, new weights and z-scores have been calculated for all households and individuals, where applicable.
Changes in Wave 1 data from Version 4.0 to 4.1:
1. Weights
Changes to the data to cater for a polygamist mandated a change to the weights.
2. Z-Scores
A mistake in the coding file for z-scores found after the release of version 4 has been corrected.
Version 5 of the NIDS Wave 1 (2008) data was received on 22 August 2013. The following changes were made to Version 4.1 of the National Income Dynamics Study (NIDS) Wave 1 dataset to produce Version 5:
Data Corrections in Version 5
Discrepancies in birth history, and parent vital status (mother/father alive) data were corrected with data from call-backs to households.
Duplicate households (resulting from interviews with the same respondents in more than one household) were identified during Wave 3 fieldwork and corrected. This has resulted in a change in the number of individuals and households in the dataset.
PIDs were created for non-resident household members in the Wave 1 HouseholdRoster file. 732 non-residents were matched to respondents in Wave 2 or 3.
Documents Renamed in Version 5
The Wave 1 household questionnaire file was renamed from HouseholdQ to HHQuestionnaire for consistency with Waves 2 and 3.
New Variables in Version 5
Indderived FILE:
w1_best_mthpid
w1 _best_fthpid
These identify mothers and fathers in the NIDS panel even when they were not co-resident with their children or had died.
w1_h_preflng_o
w1_best_gen
HouseholdRoster FILE:
w1_r_res
Renamed Variables in Version 5:
Variables have been renamed in the Child file to ensure consistency in the variable names across files. Please see the User Manual for a list of renamed variables
Dropped Variables in Version 5:
HouseholdRoster FILE
wx_r_age (use best_age in the individual derived file)
Most of the variables dropped were empty variables. Please see the User Manual for a list of the dropped variables
New Weights in Version 5:
All weights were recalculated in version 5. Please see the NIDS Wave 3 User Manual for an explanation of how the weights were calculated the relationship between the different weights.
Changes in Version 5.1
Admin data
Admin data has been added to the regular wave specific pack. Previously this was a separate item to download via the DataFirst catalogue. We hope that this convenience will enrich users' experience of developing research from this ever growing resource. The publically available data matches the names of schools as collected by NIDS to Department of Basic Education's Ordinary School's Master List. Only a limited number of variables are made publically available to protect the identities of NIDS respondents. A secure data facility is provided where researchers can match their own data sources based on EMIS numbers to the matched schools. See <http://www.nids.uct.ac.za/nids-data/secure-data> for further details.
Renamed variables
The variable w1_pi_hhimprent was incorrectly named in the hhderived file in Version5.0. It has been renamed back to w1_hhimprent_inc.
Birth History changes
There were a few changes on the pcode and pid on 8 of the individuals listed on the BH section. An incorrect pcode and pid had been wrongly allocated to these individuals. These have since been corrected.
Pcode changes
2 individuals had been incorrectly assigned a pcode of 44 which is invalid. This error has been fixed on the HouseholdRoster file by assigning the correct pcode for the two respondents.
Svyset
Through interaction with our users it was brought to our attention that the svyset command in STATA was retaining settings. We have subsequently removed these settings from all data sets.
Changes in V5.2 (February 2014)
Weights
NIDS datasets have been reweighted to take into account the Census 2011 geographic data. The change in the weights have also impacted slightly on the w1_hhquint variable as it represents the weighted household quintile.
Renamed Variables
Previous geographic variables have been given the suffix '2001' to distinguish them from the new geographic variables. The following variables were affected:
w1_hhprov became w1_hhprov2001
w1_hhgeo becamew1_hhgeo2001
w1_hhdcbecame w1_hhdc2001
w1_hhmp became w1_hhmp2001
w1_hheabecame w1_hhea2001
w1_mapped_prov*became w1_mapped_prov2001*
w1_mapped_dc* became w1_mapped_dc2001*
w1_mapped_mp* became w1_mapped_mp2001*
w1_mapped_geo* became w1_mapped_geo2001*
w1_mapped_ea* became w1_mapped_ea2001*
*Secure dataset variables
New Variables
Census 2011 Geographic Variables have been brought into the NIDS dataset. The new variables are:
New Variable Name
w1_hhprov2011 w1_hhgeo2011 w1_hhdc2011 w1_hhmdbdc2011 w1_hhmp2011 w1_hhea2011 w1_mapped_prov2011* w1_mapped_dc2011* w1_mapped_mdbdc2011* w1_mapped_mp2011* w1_mapped_geo2011* w1_mapped_eatype2011* w1_mapped_ea2011*
*Secure dataset variables
More detail about this change can be found in the document detailing the Inclusion of Census 2011 data in NIDS.
Changes in version 5.3
Version 5.3 of NIDS Wave 1 2008 had minor changes to some variable lables.
CHANGES IN VERSION 6
New Variables/Info:
Migration Variables
NIDS collects data on locations in which respondents have lived in the past (Questions b10 - b16 in the Adult). This migration data was previously coded using 2001 Census data to district municipality level (DC). In the latest release this migration data is are now coded to both the 2001 and 2011 Census data and the new release has both versions of the district municipality codes. New variables for migration have the suffix dc_2001 and dc_2011 for descriptions coded to the 2001 and 2011 Census data respectively.
Birth Histories
NIDS has tried to identify and match all the children across Wave 1 - Wave 4 in the Birth History (BH) section. Respondents were contacted for further information, and changes made based on new information received. An additional gain from this exercise is that each child in the BH section now has a PID to identify them.
Negative and Positive events Variables
The following “other” variables in the household questionnaire are now available in the HHQuestionnaire data file: w1_h_nego_o, w1_h_poso1_o & w1_h_poso2_o.
Employment Codes
Employment codes for Wave 1 data were previously created using the South African Standard Classification of Occupations (SASCO) codes, whereas subsequent waves used the International Standard Classification of Occupations (ISCO) codes. In order to make Wave 1 consistent with the other waves, all Wave 1 employment descriptions have now been coded using the ISCO codes. The variable names for the codes created have been changed to indicate this. Disaggregated ISCO codes up to the 5-digit level are available in the Secure (restricted access) version of the data, while the one-digit level codes are included in the Public Release data.
Some respondents in the Adult dataset had answered that they had other self-employment activities. However, occupational codes for "other self-employment activities" did not exist in the data. This new variable has been added: w1_a_emsothatc_isco_c, which has the occupational codes of respondents with other self-employment activities.
Police District data
Police district data has now been included as part of the Admin data file. Variables include distance to the nearest police station and distance to the police station in the district in which the household is located. Only categorical distances have been included in the public release version of the data. Actual distances can be found in the secure (restricted-access) version of the data.
Interviewer Evaluation
Substantive cleaning was done on the variables indicating which household member helped to complete the questionnaire. This resulted in a fourth person being created for the Proxy questionnaire. The variable added is w1_p_intresppid4.
Parental data
An exercise to reduce inconsistences in the parental information was carried out for all individuals across all waves. Cases with problems were identified by comparing parental information across waves. Information obtained from contacting respondents was used to correct inconsistent parental data, where possible.
Pcode variables
The pcode variables have been dropped from Wave 1 data. This was done for the following reasons:
All non-resident individuals now have a pid, thus the pcode becomes a duplicate identifier
The task of cleaning the identifiers was becoming cumbersome, as every pid adjustment required a pcode adjustment
The pcode did not exist in any of the later datasets. We therefore dropped these variable in order to have consistency in the panel.
Re-named variables in version 6 are listed in the document on changes in this version.
CHANGES IN VERSION 6.1
Version 6.1 has changes to the weight variables, w1_pweight in the indderived data file and the w1_wgt in the hhderived data file. The weight variables were changed because:
1. Panel weights were missing for some babies born to CSM mothers after Wave 1 (2008)
2. The weight for one respondent was missing
3. This version includes a syntax file to correct an error concerning the w*_a_unemwnt (number of years wanting work with no success) variable in the Adult data file. This variable was inconsistently re-named across the panel. This variable name will be corrected in the next NIDS data release.
CHANGES IN VERSION 7.0.0
Version 7.0.0 of NIDS wave 1 2008 includes changes to the amount of individuals and households in each data file, largely driven by previously incorrect classification of TSM/CSM status, duplicate interviews and additional baby CSMs not captured in a previous version of this wave. Version 7.0.0 also contains new and renamed variables, and there are changes to the survey weights. For details on these changes please see the document Wave 1 Changes between V6.1 and V7.0.0 provided with the data.