Version Notes
VERSION NOTES
Version 4 of the NIDS Wave 1 secure data, produced 20 February 2012, had the following changes from version 3:
1. Sample size (one household dropped).
Household 109011 has been dropped from the sample. This household was interviewed twice under phase 1 (with hhid 109009 ) and under phase 2 ( with hhid 109011). The latter interview has been excluded from the sample. As a result there is 1 fewer household in the HouseholdQ file. The two individuals (an adult and a proxy) in this household have been dropped from the HouseholdRoster file and the individual level files.
2. Duplicate record for polygamous relationship in HouseholdRoster file
A duplicate record has been created in the HouseholdRoster file for pid 316891 with different hhid's. This individual is resident in two households (106506 and 106744). The implication of this is that individuals in the HouseholdRoster file are now uniquely identified by hhid and pid, rather than pid only. This will be useful to remember when it comes to merging the different STATA files together.
Note that because of the polygamist's pid appearing in two different households, merges should take place on both hhid and pid in Version 4.0 of the data.
3. New Variables added
The following variables have been added to the public release dataset:
a. Interviewer Evaluation variables have been included in the Adult, Child, Proxy and HouseholdQ files.
b. New pid variables have been added for every question that asks for the pcode of the respondent.
c. Date of interview day has been included in all datasets. Users have found this variable useful, hence its inclusion.
4. Variables dropped
a.The household questionnaire variable, hhqi, has been dropped from the wave1 dataset, to avoid confusion with the wave 2 hhid variable in terms of content. The hhqi variable also no longer appears in the wave 2 dataset.
b. Hhderived weight variables have been dropped
5. Variable names updated
a. Year of interview has been included in Education variable names in the Adult, Child and Proxy files
b. The hhid variable has been renamed to w1_hhid in all datasets
6. Update of Key variables, Weights and Z-scores
In the process of undertaking Wave 2 fieldwork, key variables (dob, age, gender, education and race) have been updated as better responses from respondents were received in cases where there had been non responses or errors in wave 1 data. As a result of changes in these key variables and samples sizes of the dataset, new weights and z-scores have been calculated for all households and individuals, where applicable.
Changes in Wave 1 data from Version 4.0 to 4.1:
1. Weights
Changes to the data to cater for a polygamist mandated a change to the weights.
2. Z-Scores
A mistake in the coding file for z-scores found after the release of version 4 has been corrected.
Version 5 of the NIDS Wave 1 (2008) data was received on 22 August 2013. The following changes were made to Version 4.1 of the National Income Dynamics Study (NIDS) Wave 1 dataset to produce Version 5:
Data Corrections in Version 5
Discrepancies in birth history, and parent vital status (mother/father alive) data were corrected with data from call-backs to households.
Duplicate households (resulting from interviews with the same respondents in more than one household) were identified during Wave 3 fieldwork and corrected. This has resulted in a change in the number of individuals and households in the dataset.
PIDs were created for non-resident household members in the Wave 1 HouseholdRoster file. 732 non-residents were matched to respondents in Wave 2 or 3.
Documents Renamed in Version 5
The Wave 1 household questionnaire file was renamed from HouseholdQ to HHQuestionnaire for consistency with Waves 2 and 3.
New Variables in Version 5
Indderived FILE:
w1_best_mthpid
w1 _best_fthpid
These identify mothers and fathers in the NIDS panel even when they were not co-resident with their children or had died.
w1_h_preflng_o
w1_best_gen
HouseholdRoster FILE:
w1_r_res
Renamed Variables in Version 5:
Variables have been renamed in the Child file to ensure consistency in the variable names across files. Please see the User Manual for a list of renamed variables
Dropped Variables in Version 5:
HouseholdRoster FILE
wx_r_age (use best_age in the individual derived file)
Most of the variables dropped were empty variables. Please see the User Manual for a list of the dropped variables
New Weights in Version 5:
All weights were recalculated in version 5. Please see the NIDS Wave 3 User Manual for an explanation of how the weights were calculated the relationship between the different weights.
Changes in Version 5.1
Admin data
Admin data has been added to the regular wave specific pack. Previously this was a separate item to download via the DataFirst catalogue. We hope that this convenience will enrich users' experience of developing research from this ever growing resource. The publically available data matches the names of schools as collected by NIDS to Department of Basic Education's Ordinary School's Master List. Only a limited number of variables are made publically available to protect the identities of NIDS respondents. A secure data facility is provided where researchers can match their own data sources based on EMIS numbers to the matched schools. See <http://www.nids.uct.ac.za/nids-data/secure-data> for further details.
Renamed variables
The variable w1_pi_hhimprent was incorrectly named in the hhderived file in Version5.0. It has been renamed back to w1_hhimprent_inc.
Birth History changes
There were a few changes on the pcode and pid on 8 of the individuals listed on the BH section. An incorrect pcode and pid had been wrongly allocated to these individuals. These have since been corrected.
Pcode changes
2 individuals had been incorrectly assigned a pcode of 44 which is invalid. This error has been fixed on the HouseholdRoster file by assigning the correct pcode for the two respondents.
Svyset
Through interaction with our users it was brought to our attention that the svyset command in STATA was retaining settings. We have subsequently removed these settings from all data sets.
Changes in V5.2 (February 2014)
Weights
NIDS datasets have been reweighted to take into account the Census 2011 geographic data. The change in the weights have also impacted slightly on the w1_hhquint variable as it represents the weighted household quintile.
Renamed Variables
Previous geographic variables have been given the suffix '2001' to distinguish them from the new geographic variables. The following variables were affected:
w1_hhprov became w1_hhprov2001
w1_hhgeo became w1_hhgeo2001
w1_hhdc became w1_hhdc2001
w1_hhmp became w1_hhmp2001
w1_hhea became w1_hhea2001
w1_mapped_prov* became w1_mapped_prov2001*
w1_mapped_dc* became w1_mapped_dc2001*
w1_mapped_mp* became w1_mapped_mp2001*
w1_mapped_geo* became w1_mapped_geo2001*
w1_mapped_ea* became w1_mapped_ea2001*
*Secure dataset variables
New Variables
Census 2011 Geographic Variables have been brought into the NIDS dataset. The new variables are:
New Variable Name
w1_hhprov2011 w1_hhgeo2011 w1_hhdc2011 w1_hhmdbdc2011 w1_hhmp2011 w1_hhea2011 w1_mapped_prov2011* w1_mapped_dc2011* w1_mapped_mdbdc2011* w1_mapped_mp2011* w1_mapped_geo2011* w1_mapped_eatype2011* w1_mapped_ea2011*
*Secure dataset variables
More detail about this change can be found in the document detailing the Inclusion of Census 2011 data in NIDS.
Changes in version 5.3
Version 5.3 of NIDS Wave 1 2008 had minor changes to some variable lables.
CHANGES IN VERSION 6
New Variables/Info:
Migration Variables
NIDS collects data on locations in which respondents have lived in the past (Questions b10 - b16 in the Adult). This migration data was previously coded using 2001 Census data to district municipality level (DC). In the latest release this migration data is are now coded to both the 2001 and 2011 Census data and the new release has both versions of the district municipality codes. New variables for migration have the suffix dc_2001 and dc_2011 for descriptions coded to the 2001 and 2011 Census data respectively.
Birth Histories
NIDS has tried to identify and match all the children across Wave 1 - Wave 4 in the Birth History (BH) section. Respondents were contacted for further information, and changes made based on new information received. An additional gain from this exercise is that each child in the BH section now has a PID to identify them.
Negative and Positive events Variables
The following “other” variables in the household questionnaire are now available in the HHQuestionnaire data file: w1_h_nego_o, w1_h_poso1_o & w1_h_poso2_o.
Employment Codes
Employment codes for Wave 1 data were previously created using the South African Standard Classification of Occupations (SASCO) codes, whereas subsequent waves used the International Standard Classification of Occupations (ISCO) codes. In order to make Wave 1 consistent with the other waves, all Wave 1 employment descriptions have now been coded using the ISCO codes. The variable names for the codes created have been changed to indicate this. Disaggregated ISCO codes up to the 5-digit level are available in the Secure (restricted access) version of the data, while the one-digit level codes are included in the Public Release data.
Some respondents in the Adult dataset had answered that they had other self-employment activities. However, occupational codes for "other self-employment activities" did not exist in the data. This new variable has been added: w1_a_emsothatc_isco_c, which has the occupational codes of respondents with other self-employment activities.
Police District data
Police district data has now been included as part of the Admin data file. Variables include distance to the nearest police station and distance to the police station in the district in which the household is located. Only categorical distances have been included in the public release version of the data. Actual distances can be found in the secure (restricted-access) version of the data.
Interviewer Evaluation
Substantive cleaning was done on the variables indicating which household member helped to complete the questionnaire. This resulted in a fourth person being created for the Proxy questionnaire. The variable added is w1_p_intresppid4.
Parental data
An exercise to reduce inconsistences in the parental information was carried out for all individuals across all waves. Cases with problems were identified by comparing parental information across waves. Information obtained from contacting respondents was used to correct inconsistent parental data, where possible.
Pcode variables
The pcode variables have been dropped from Wave 1 data. This was done for the following reasons:
All non-resident individuals now have a pid, thus the pcode becomes a duplicate identifier
The task of cleaning the identifiers was becoming cumbersome, as every pid adjustment required a pcode adjustment
The pcode did not exist in any of the later datasets. We therefore dropped these variable in order to have consistency in the panel.
Variables Re-named
Table 2 below shows all the variables that have been renamed in V6.0 data.
Quest. Section Old Variable name New Variable Name
Adult Demographics w1_a_movy w1_a_moveyr
Adult Demographics w1_a_brndc w1_a_brndc_2001
Adult Demographics w1_a_lvbfdc w1_a_lvbfdc_2001
Adult Demographics w1_a_lv94dc w1_a_lv94dc_2001
Adult Demographics w1_a_lv06dc w1_a_lv06dc_2001
Adult Labour Market Participation w1_a_em1occ_c w1_a_em1occ_isco_c
Adult Labour Market Participation w1_a_em2occ_c w1_a_em2occ_isco_c
Adult Labour Market Participation w1_a_emsatc_c w1_a_emsatc_isco_c
Adult Labour Market Participation w1_a_emsoth_c w1_a_emsothatc_isco_c
Adult Labour Market Participation w1_a_emctype_c w1_a_emctype_isco_c
Adult Labour Market Participation w1_a_emhtsk_c w1_a_emhtsk_isco_c
Adult Labour Market Participation w1_a_unemtyp_c w1_a_unemtyp_isco_c
Adult Parents' vital status w1_a_mthwrk_c w1_a_mthwrk_isco_c
Adult Parents' vital status w1_a_fthwrk_c w1_a_fthwrk_isco_c
Adult Labour Market Participation w1_a_eminc w1_a_emcinc
Adult Labour Market Participation w1_a_em1inc_s w1_a_em1inc_sh
Adult Labour Market Participation w1_a_eminc_sh w1_a_emcinc_sh
Adult Education w1_a_ed07payr1 w1_a_ed07paypr1
Adult Education w1_a_ed07payr2 w1_a_ed07paypr2
Adult Education w1_a_ed07payr3 w1_a_ed07paypr3
Child Demographics w1_c_movy w1_c_moveyr
Child Demographics w1_c_brndc w1_c_brndc_2001
Child Demographics w1_c_lv06dc w1_c_lv06dc_2001
Child Demographics w1_c_lvbfdc w1_c_lvbfdc_2001
Child Parents' vital status w1_c_mthwrk_c w1_c_mthwrk_isco_c
Child Parents' vital status w1_c_fthwrk_c w1_c_fthwrk_isco_c
Child Education w1_c_ed07wdexp w1_c_ed07wdex
Child Education w1_c_fthedlev w1_c_fthtert
Child Education w1_c_mthedlev w1_c_mthtert
Child Education w1_c_ed07payr1 w1_c_ed07paypr1
Child Education w1_c_ed07payr2 w1_c_ed07paypr2
Child Education w1_c_ed07payr3 w1_c_ed07paypr3
Child Parents and family support w1_c_fththa w1_c_fthdtha
Child Grants w1_c_grcurecr w1_c_grcurecrel
Child Child's Health w1_c_hlthdes w1_c_hldes
Child Parents' vital status w1_c_mthtrt w1_c_mthtertyn
Child Parents' vital status w1_c_fthtrt w1_c_fthtertyn
hhderived w1_hhcluster w1_cluster
hhderived - w1_hhdc2001 w1_dc2001
hhderived - w1_hhdc2011 w1_dc2011
hhderived - w1_hhgeo2001 w1_geo2001
hhderived - w1_hhgeo2011 w1_geo2011
hhderived - w1_hhprov2001 w1_prov2001
hhderived - w1_hhprov2011 w1_prov2011
hhderived - w1_hhmdbdc2011 w1_mdbdc2011
HHQ A w1_h_pcode_pid w1_h_respondent
HHQ Negative Events w1_h_negdthfrin w1_h_negdthfrinc
HHQ Agriculture w1_h_*prd w1_h_*prdss
Proxy Demographics w1_p_movy w1_p_moveyr
Proxy Demographics w1_p_brndc w1_p_brndc_2001
Proxy Demographics w1_p_lv06dc w1_p_lv06dc_2001
Proxy Demographics w1_p_lv94dc w1_p_lv06dc_2011
Proxy Demographics w1_p_lvbfdc w1_p_lv94dc_2001
Proxy Labour Market Participation w1_p_emp w1_p_emactcur_u
Proxy Labour Market Participation w1_p_empinc w1_p_em1inc_sh
Proxy Labour Market Participation w1_p_empocc_c w1_p_em1occ_isco_c
Proxy Labour Market Participation w1_p_empprod_c w1_p_em1prod_c
CHANGES IN VERSION 6.1
Version 6.1 has changes to the weight variables, w1_pweight in the indderived data file and the w1_wgt in the hhderived data file. The weight variables were changed because:
1. Panel weights were missing for some babies born to CSM mothers after Wave 1 (2008)
2. The weight for one respondent was missing
CHANGES IN VERSION 7.0.0
Version 7.0.0 of NIDS wave 1 2008 includes changes to the amount of individuals and households in each data file, largely driven by previously incorrect classification of TSM/CSM status, duplicate interviews and additional baby CSMs not captured in a previous version of this wave. Version 7.0.0 also contains new and renamed variables, and there are changes to the survey weights. For details on these changes please see the document Wave 1 Changes between V6.1 and V7.0.0 provided with the data.