遇见数据集

Household Income and Living Survey Data Collection

收藏
DataInfoPlus2026-07-17 收录
官方服务:

资源简介:

Target population and survey scope The survey scope defines the target population for the survey. This then defines the benchmarks that are used (the size of the target population is equal to the total of the benchmarks). The target population is all usually resident individuals of private dwellings in urban and rural areas in North Island, South Island, and Waiheke Island of Aotearoa New Zealand. The following people are out of scope: overseas visitors who have been or expect to be in the country for less than 12 months residents of non-private dwellings such as hotels, motels, hostels, and boarding houses long-term residents of institutions such as hospitals and/or prisons persons in homes for the aged (including rest homes) where there are communal cooking facilities (where long-term is defined as more than six weeks) members of permanent armed forces who live in non-private dwellings, such as barracks members of armed forces serving overseas international diplomats and their families usual residents of offshore islands (except for Waiheke Island). Survey population The survey population is the target population with some exclusions due to practical survey difficulties. The following are excluded: usual residents temporarily overseas who do not return within the survey period usual residents temporarily staying elsewhere in Aotearoa New Zealand who do not return to the selected household within the survey period people residing at a wharf or landing place (for example, people on ships). Children over the age of 15 who are away at boarding school who would be surveyed if they lived at home are included as part of their parents’/caregivers’ household. Sample design information HILS uses a stratified, multi-stage, cluster design. Primary sampling units (PSUs) – a geographic unit – are selected from the household survey frame, then dwellings within PSUs, then eligible persons within selected dwellings. We usually plan to use selected PSUs for several years, selecting different dwellings each year. Using the same set of PSUs, or having overlap with a previous year, provides some stability in the sample characteristics between years and is an efficient use of the surveyed PSUs. In the case of the HILS as with HES, a selected PSU can supply three years of survey sample. The set of PSUs selected for the survey expansion in 2018 were used from 2018/2019 to 2020/2021. A new sample of PSUs was selected for the now discontinued Living in Aotearoa survey in 2021, using an updated sample design based on that introduced for HES in 2018/2019. The information on the frame used in the design was updated using 2018 Census data. A subsample of 4,368 PSUs were selected for Living in Aotearoa from the household survey frame. From this sample of 4,368 PSUs, sub-samples of 2,500 were selected for each of the 2021/2022, 2022/2023, and 2023/2024 HES years. These all had some overlap with each other. After Living in Aotearoa was discontinued, the 2024/2025 HILS also selected 2,220 PSUs from the wider Living in Aotearoa sample, which are all common to the 2023/2024 HES. For the 2025/2026 HILS, collected from July 2025 to June 2026, PSUs were selected from a frame that has been updated using 2023 Census data. This new sample used the updated design previously used for HES and Living in Aotearoa. PSUs are stratified according to census information for each PSU on the household frame. Stratification primarily helps to ensure that different subgroups are represented in the sample (for example, regional council areas, NZ Deprivation (NZDep) index areas), as well as to manage the sampling rate in more costly strata to reduce total survey cost, and to reduce the sampling variance of the estimates. The sample was stratified by: region – 12 geographical strata based on regional council areas (West Coast, Tasman, Nelson, and Marlborough were collapsed into one region, as were Gisborne and Hawke’s Bay, to create larger regions) urban/rural – we used an urban/rural sampling ratio of 1.4:1 to control collection costs for Stats NZ, as rural PSUs are more expensive than urban PSUs to survey NZDep2018 Index – the inclusion of NZDep2018 in the stratification ensures a good spread of areas by socio-economic status estimated child poverty indicators from census data (at PSU level) – this ensures PSUs with high numbers of lower socio-economic children are represented household income – defined as total gross income from the 2018 Census. Next, we select dwellings within selected PSUs. On average 11.4 dwellings are selected within each PSU, as we aim to achieve around eight households per PSU in our final sample. To do this, dwellings within the PSU are allocated systematically to one of six groups (called panels). Each annual survey uses two of these panels – one for the main survey and one for increasing the sample size of Māori (see Oversampling Māori). Finally, we interview all individuals aged 15 years and over within each selected dwelling. Selections are distributed across the 12-month survey period so that survey results are representative of income patterns across the year. Oversampling Māori For the core HILS, we oversample Māori to increase the likelihood of achieving a higher number of Māori in our sample than we would by chance. Since 2021/2022, we have used Māori descent information from the 2018 Census to identify if a dwelling has a Māori household member. It enables us to identify dwellings likely to contain at least one person identifying as Māori and then sample them at a higher rate. From 2018/2019 to 2020/2021, the electoral roll, rather than census data, was used for this purpose. Within each PSU, dwellings with at least one Māori household member have a higher chance of being selected into the sample. The resulting differential selection probabilities are adjusted for during weighting. Analysis comparing the use of the electoral roll and census data showed that either method would pick up a similar number of households. Sample size When HES was re-designed to produce child poverty statistics that meet Stats NZ’s responsibility under the Act, accuracy objectives were set to limit the size of the sample error associated with annual changes. The HES target sample size was set at 20,000 responding households. This limited the size of the sample error on national-level child poverty estimates to about 1.5 percentage points across most of the poverty measures, and 3.9 percentage points for tamariki Māori (although there was some fluctuation in the expected sample error size across different poverty measures). This design allowed for the survey to detect real-world annual changes on the magnitude of 1.5 percentage points and 3.9 percentage points or more (for the total population and tamariki Māori, respectively). Collection challenges, driven by factors, including COVID-19 and increasing collection costs, meant that the target sample size for HES was achieved only in the 2018/2019 collection year. While obtaining a large enough sample is important, it is also important to consistently achieve the target sample size during collection, because the sample is carefully designed to provide a representative picture of the population, within a given level of precision. When targets are not met by a large margin, there is heightened risk of impacts to data quality, including a risk of bias. To ensure the ongoing sustainability of producing high-quality statistics, the sample size design for HILS was re-assessed ahead of the first HILS collection in July 2024. This assessment determined that a sample size of 17,000 households would achieve the balance of accuracy and sustainability into the future, enabling the continued production of high-quality child poverty statistics that meet Stats NZ’s responsibilities under the Act. The main impact of the reduced sample size is that the detection threshold for identifying annual change has increased by about 0.1 to 0.2 percentage points on average, for the total child population, and about 0.3 percentage points for tamariki Māori (compared to if 20,000 households were interviewed). As with HES, we encourage users to focus on estimates of change over longer time frames, such as the three-year target periods. With 2,220 PSUs, and just over 11.4 households selected per PSU, we get a total selected sample of approximately 25,000 households. We assume that at least 68 percent of households will provide a full response, leaving a final achieved sample of at least 17,000 households. (The target sample sizes for the expenditure and net worth components, which are sub-samples of the core HILS, are 5,500 and 8,500 households, respectively, and these are unchanged from HES.) The final achieved sample in the 2024/2025 HILS included approximately 17,892 households. Reliability of survey estimates Two types of error are possible in estimates based on a sample survey – sample error and non-sample error. Sample error is a measure of the variability that occurs by chance because a sample, rather than an entire population, is surveyed. We can calculate the level of uncertainty around a survey estimate by exploring how that estimate would change if we were to draw many survey samples for the same period instead of just one. Sample errors are also calculated for estimates of change by considering the uncertainty associated with the estimate at each of the two time points of interest. This allows us to define a range around the standalone estimate or estimate of change, and to state how likely it is that the real value that the survey is trying to measure lies within that range. These ranges are referred to as confidence intervals and are typically set up so that we can be 95 percent sure that the true value lies within the range – in which case this range is referred to as a 95 percent confidence interval. Confidence intervals are used as a guide to the size of the sample error. A wider confidence interval indicates a greater uncertainty around the estimate. Generally, a smaller sample size will lead to estimates that have a wider confidence interval than estimates from larger sample sizes. This is because a smaller sample is less likely than a larger sample to reflect the characteristics of the total population, and therefore there will be more uncertainty around the estimate derived from the sample. The 95 percent confidence interval is used in HILS reporting and is calculated as the estimate plus or minus the sample error. We calculate sample errors using the jackknife method, which is based on the variation between estimates of different sub-samples taken from the whole sample. Sample errors can be used to identify changes in the data that are due to real-world effects and are unlikely to have occurred by chance due to a particular sample being chosen. If an observed annual change is larger than the associated sample error on the change, this change is unlikely to be the result of chance and is therefore considered ‘statistically significant’. With an intended achieved sample size of 17, 000 households, it is expected that: sample errors (95 percent confidence intervals) for the annual change in rates for the nine child poverty measures will be approximately 1.7 percentage points or less sample errors (95 percent confidence intervals) for the annual change in rates for Māori children will be approximately 4.2 percentage points or less. Sample errors can also scale with size of the group estimated. One main exception to the precision targets above is for measure (f), where there tend to be larger sample errors because the proportion in poverty is higher than for other measures. For all children, the sample errors associated with annual change from 2023/2024 to 2024/2025 ranged from 1.1 to 2.0 across the nine poverty measures. For Māori children, they ranged from 2.8 to 4.4. Non-sample error can occur in any survey, whether the estimates are derived from a sample or a census. Sources of non-sample error include non-response, errors in respondents’ reporting or interviewers’ recording of answers, and errors in data processing. Every effort is made to minimise non-sample error by careful design and testing, training of survey interviewers, and editing and quality control procedures during data processing. Any remaining error is very difficult to identify and quantify. Non-response can affect the reliability of results and introduce bias if the people who do not respond differ systematically in some important characteristic from those who do respond. For example, if the response rate is low among people with low income, not only can we be less confident in the income estimates for this group, but national estimates will also be biased towards higher incomes. We employ additional effort in the field to achieve as high a response rate as possible from low socio-economic groups and from different regions. Our weighting methodology is also designed to mitigate the impact of lower response rates from certain subgroups of the population (that is, by adjusting the weights upwards). However, some bias will remain if the missing respondents have substantially different income to those who do respond. Imputation We use imputation in the core HILS to replace missing data for households with partially completed surveys (item non-response), as well as for non-responding individuals (unit non-response) residing in otherwise fully responding households. Households are defined as fully responding when the HQ is complete, as well as the PQ of the nominated best person to answer financial questions. The demographic information collected allows us to impute the record of the non-responding household member by linking to the IDI or via imputation software (described below). Rather than discarding incomplete records, these methods allow us to make the best use of the data collected in HILS. We use imputation software when IDI information is unavailable, for the following variables: Income: employment earnings and government transfers where a respondent has not been linked to the IDI self-employment income where respondent is known to have such but has not provided a value investment income where respondent is known to have such but has not provided a value. Housing costs: local and regional authority property rates for primary property. Person demographics: age gender sex ethnicity disability status for people over the age of 2 years. We use the Canadian Census Editing and Imputation System (CANCEIS) software developed by Statistics Canada to perform deterministic and nearest neighbour donor imputation for the variables above. The nearest-neighbour imputation method identifies a donor (respondent) ‘nearest’ to the recipient (the non-respondent). ‘Nearest’ is defined using a distance function based on other known (explanatory or matching) variables correlated with the missing values. The distance function involves assigning a weight to a matching variable. A larger weight is given to a variable if it is believed to be more accurate and is a good predictor for the missing variable(s). Thus, to impute a missing income value for a recipient, CANCEIS will find suitable donors based on matching variables and their weights. Consistency edits or outlier edits are included in CANCEIS to minimise selection of donors that: create implausible combinations of data for the recipient disproportionately increase the number of times an outlying donor is used – for example, for people aged 15 to 24 years with a missing marital status, there is a limit on imputing their status with widow/widower have imputed data for the matching variables. We used machine learning – specifically, classification and regression trees and random forests – to efficiently impute using the nearest-neighbour imputation method, by determining the matching variables and corresponding weights. Evaluations carried out for the 2018/2019 HES showed our imputation methodology to be effective, showing high predictive and distributional accuracy. In the 2024/2025 HILS, 36 percent of households and 22 percent of individuals have had one or more imputed value. These imputation rates are higher than what was typically observed in HES. This difference reflects the introduction of the proxy module in HILS, which captures limited demographic information for IDI linking purposes. Much of this difference was driven by an increase in the proportion of records that had disability status imputed, because disability questions are not asked in the proxy module (whereas in HES, proxy respondents would have gone through the full questionnaire and answered disability, and other questions, on behalf of the proxy). For the 2024/2025 HILS, 14 percent of people aged 2+ had disability status imputed, compared to 7 percent in the 2023/2024 HES. Note that it is possible for people aged 15+ to have zero personal income, even after imputation. This can occur because they are successfully linked to the IDI and are recorded as having zero income in admin data as well as in the survey. Those with zero income can also become donors for imputation, leading others to also have zero incomes. Weighting Weighting is used to estimate the population from the sample. A weight is attached to each unit in the sample that indicates the number of households and people it represents in the final population estimate. Weighting ensures that estimates reflect the sample design and align with current population estimates. Our weighting methodology is also designed to mitigate the impact of lower response rates from certain subgroups of the population (that is, by adjusting the weights upwards). However, some bias will remain if the missing respondents have substantially different income to those who do respond. For HILS, deriving weights is a multi-stage process. First, we calculate a household’s initial weight. This depends on the sample design and equals the inverse of the household’s selection probability (which itself depends on the selection probability of the PSU from the household survey frame, as well as the selection probability of the dwelling within the PSU). For 2024/2025, the selection probability for the dwelling was calculated as the inverse of the number of panels in the PSU (for example, 1/6). This method is used in other Stats NZ surveys and aligns with common practice. Second, we adjust the initial weights to account for unit non-response. Non-responding households are given weights of zero, while the initial weights of responding households are scaled by a rate-up factor based on the inverse of the weighted response rate of households. This is done in weighting classes formed by cross-classifying variables that are correlated with likelihood to respond. The weighting classes used for the 2024/2025 HILS were: region, NZDep2018, ethnic densities, urban/rural, and interview quarter. This step creates adjusted response weights from the initial weights. Finally, we calibrate the adjusted response weights so that estimates reflect expected population totals or benchmarks. Calibration adjusts for under-coverage of the target population. We use a form of calibration called integrated weighting to ensure that all individuals in the same household are given the same weight and that household statistics derived from person-level data match the same statistics calculated directly from household-level data. For 2024/2025, we used benchmarks based on the estimated resident population (ERP) and administrative data (admin data) on income and benefit receipt available in the Integrated Data Infrastructure (IDI). The ERP for a particular year uses census information adjusted for census coverage, births, deaths, and net migration since the most recent census (the base year). The 2023 Census was the base year for the ERP used in the 2024/2025 HILS. We calibrated the HILS income distribution of adults to the income distribution of adults available in the IDI, and then calibrated these (adjusted) weights to the other benchmarks (that is, ERP benchmarks and the number of people in the IDI who received any government benefit). The benchmark variables/categories used in the calibration process are listed below. Benchmarks based on the ERP are: children – three 5-year age groups: 0‒4, 5‒9, 10‒14 years sex by age groups – males and females by 14 age groups: 15‒17, 18‒19, 20‒24, 25‒29, 30‒34, 35‒39, 40‒44, 45‒49, 50‒54, 55‒59, 60‒64, 65‒69, 70‒74, 75+ years region – 12 regions: Northland, Auckland, Waikato, Bay of Plenty, Gisborne-Hawke’s Bay, Taranaki, Manawatū-Whanganui, Wellington, West Coast-Tasman-Nelson-Marlborough, Canterbury, Otago, and Southland Māori adults by age – two age groups for Māori: 15‒29, 30+ years households by region and household type – 12 regions for households with two adults or without two adults, separately. Benchmarks from admin data are: people who received any government benefit, excluding New Zealand Superannuation and Veterans’ Pension the income distribution of adults – income deciles using total income (sum of income from all regular income sources) at the individual level. Updates to previous weights The weights for HES from 2018/2019 to 2023/2024 used the 2018-base ERP when statistics for those years were first produced. These weights have now been updated using the 2023-base ERP. On 12 February 2026, we published revised household income and housing cost statistics for the years ended June 2019 to 2024, as part of the population rebase. Household Economic Survey population rebase: Year ended June 2019 to 2024 provides further information on the population rebase and the impact on estimates. Data collection methodology Stats NZ data collection specialists visit selected households and conduct computer-assisted interviews with each eligible household member. HILS is optimised for computer-assisted in-person interviews, though for the last four years data collection specialists have also conducted computer-assisted interviews with respondents over the phone to maximise responses. In-person and phone interviewing are two examples of interview ‘mode’. We use computer-assisted interviewing software called Blaise to conduct the interviews, which guides the interviewer through the correct sequence of questions. Questions are automatically routed to ensure that respondents are only asked the questions appropriate for them. Salesforce is the centralised system we use to manage data collection, including assigning data collection specialists to households and real time monitoring of response rates. Three questionnaires are included in HILS: The household questionnaire (HQ) – collects information on household characteristics, such as household membership, relationships, and age. It is completed by one adult in the household. The personal questionnaire (PQ) – collects information from usual residents aged 15 years and over, with specific content dependent on respondent roles as determined in the HQ: For all eligible respondents, includes modules covering employment, income, and person-level material wellbeing. For the person nominated by the household as being the best person to answer household financial questions, also includes modules covering housing, business, property, and mortgage expenditure, and household-level material wellbeing. For the person(s) nominated to answer questions about children, also includes modules covering child demographics and child-level material wellbeing for the children in question. The demographic questionnaire – collects personal demographics about each usual resident aged 15 years and over. In most cases this is self-completed by the respondent on a tablet to preserve privacy, as some of the questions are sensitive, but can be interviewer administered if necessary. To ensure accurate information is collected, we aim to interview selected respondents directly. In some circumstances, however, we cannot reach all household members. Where this occurs, we may ask another person in the household to respond on behalf of the selected respondent. This is known as a proxy response. Proxy responses are used in ‘family-type’ households for those aged 15+ who would otherwise have received a PQ, and cover: those unable to be present during the interview, but who were in households where all had agreed to participate those who are away at boarding school people who are elderly, sick, or mentally incapacitated. In HILS, there is a specific questionnaire module that collects only a limited set of demographic information that enables gives us the best chance at linking the individual to admin data records (see Linking administrative data). In all proxy interviews, the interviewer must be convinced that the proxy respondent is sufficiently familiar with the selected respondent’s information and the person being proxied for must agree to the proxy. Data collection specialists use a range of strategies to obtain a good response rate from all individuals, including people who are more difficult to survey. They are trained to prepare well for each interview, to identify potential barriers at the doorstep, and to use the resources available to them to ensure interview success. They are taught to maintain a positive and neutral attitude, which bolsters respondent cooperation, compliance, and confidence in the professionalism of the interviewers. Occasionally selected respondents are unwilling or unable to complete the survey, despite an interviewer’s best efforts. As a first step, we send follow-up letters to households that we have trouble contacting. Following that, someone from our relationship management team will phone and email to try and arrange an appointment. This happens at different times and days, including weekends. Data processing methodology Once a household has been interviewed, the data is submitted to a central data store and a response logged in Salesforce. The data is then fed through various editing stages, before being loaded into the processing database. Administrative data on income Despite our best efforts to obtain accurate information about respondents’ income, relying on survey responses for data of this nature inevitably introduces uncertainty. Respondents may not remember or may fail to disclose all sources of income over the past year to the interviewer. They may provide only ‘rough estimates’, describe income after tax, or forget changes to their regular income over the year. In some cases, family members may not know the income of other family members. Similarly, benefit income is often understated, particularly when it is received for only short periods throughout the year. We combine survey data on income with admin data from the IDI, a large research database managed by Stats NZ that holds microdata about people and households. It contains full tax information related to individuals, including data provided by employers for each employee (the employee monthly schedule), self-employment income, and some investment income. We also use data provided by the Ministry of Social Development about benefits paid, including Working for Families (WFF) tax credits, and Accommodation Supplement. Since 2018/2019, we have sourced annual salary and wage, and government transfer income from admin data. In the 2018/2019 HES, we asked respondents for their income but used the admin data in the published statistics. From 2019/2020, respondents were no longer asked to provide their income amounts for these income variables. Some income sources are not currently available in the IDI, including some sources of irregular income, and non-taxable income. We collect these income variables, as well as self-employment and investment income, directly from respondents. Salary and wage income is provided on an individual’s pay day and is updated in the IDI on a quarterly basis. However, self-employment and investment income rely on individuals providing their tax returns, which may be delayed before being included in the IDI. Due to this timeliness issue, we ask respondents to provide us with information on these income sources directly. Information on income received from WFF is available from Inland Revenue and from the Ministry of Social Development and this data is used for relevant households. However, for some households this income is received annually and there can be delays in this information being incorporated in the IDI due to delays in filing tax returns. For this reason, annual income from WFF is estimated for some families and is revised in the following year when more information is available. Consistency with other periods Although we adjust survey results for various demographic variables (age, sex, and region), there can be variability in survey estimates from one survey collection period to the next. This variability is because a different group of households is selected for each survey. Suppressed estimates For confidentiality purposes, we suppress data in the released tables if a cell is based on fewer than six people or households. Any suppressed cells are identified in the tables with an ‘S’. For information on Methodology, please see Child poverty statistics: Year ended June 2025 – technical appendix en-NZ

提供机构:
Stats NZ
二维码
社区交流群
二维码
科研交流群
商业服务