Pages

Showing posts with label data. Show all posts
Showing posts with label data. Show all posts

Tuesday, February 5, 2019

Review of AEA sessions in Atlanta (Jan 4-5)

I took last month off from this blog (and most other productive activities) because I was on holiday for three weeks in the Bay Area. I hope all of you had a great holiday season 2018 with friends and family and a refreshing start to the new year 2019. The first topic I wanted to come back to is a review of the webcast sessions from the American Economic Association's annual meetings held in Atlanta from Jan 4-5, 2019. Several of the sessions are webcast here and you can access lectures on various topics including growth in the developing world, automation and the future of work, public debt, and - returning from last year with an extremely compelling panel - the gender problem in economics and what steps the profession can take to address it. In this post, I discuss two of the panels with an eye to discussing Autor's lecture on the future of work in the next post.

Growth challenges in the developing world 

The AEA convened a "World Bank economists" session consisting of three former World Bank Chief Economists (Justin Lin, Francois Bourguignon, Kaushik Basu), current Chief Economist Pinelopi Goldberg, and moderated by former Acting Chief Economist Shanta Devarajan. The purpose of the panel was to deliberate on the challenges facing the developing world. Given the very broad - arguably too broad - scope of the topic, it is natural that the panelists settled on a narrower topic over the course of the conversation: industrialization and the informality trap facing Africa.

Historically, industrialization and the rapid job creation in the formal wage sector that accompanies it have been seen as the most effective ways to raise wages and lower the poverty rate in developing countries. Lin cited historical examples of low-income countries' growth trajectories after capturing manufacturing jobs moving from the U.S. to Japan in the aftermath of WWII, Japan to Southeast Asia in 1960s and 1970s, and from Southeast Asia to China in 1980s and 1990s. Now that wages are increasing in China, many of these manufacturing jobs will be looking for a new home. How can Africa capitalize on these opportunities in coming decades was the question most of these economists were trying to answer. Chapter 2 of this policy report from the African Development Bank does a good job of summarizing these issues including evidence of what some economists call "de-industrialization" and the obstacles to small business growth. Given the demographic changes that will add 2 billion to the working age population in the African continent in this century, the creation of jobs in the formal wage sector will be important not only for economic but social and political stability. 
  1. The primary point of contention is that it is not clear that "de-industrialized" countries will capture these manufacturing opportunities without concerted policies. E.g. automation is a real threat to manufacturing jobs in certain industries and less so in others (retail incl. clothing, shoes, and furniture). Furthermore, the trade environment is rapidly changing with advanced economies looking to be less hospitable to imports from low-income countries. The second half of the panel asked panelists to comment on different ways of approaching this issue wherein I think the issue of too broad a topic came to light. I think it would have been more useful to showcase specific examples and evidence from recent research. 
  2. It wasn't discussed in the panel but it is relevant discuss the impact of a shift from self-employment and agriculture to industrial employment on working populations and whether there is desire on the part of working populations to hold these types of jobs in the first place. Specifically, J-PAL poses the issue in preface to a 2017 paper from Chris Blattman and Stefan Dercon that studied the effects of industrial employment on Ethiopian workers: "Industrial sector development to boost mass hiring is seen as important to poverty alleviation at the macroeconomic level. But how those jobs, particularly in early stages of industrial sector development, affect the workers themselves and what the workers prefer are less well-understood." The findings from this paper are summarized in this New York Times article with the bottom line being: workers are initially unaware but quickly become aware of the safety hazards and poor wages paid in sweatshop conditions leading to a high turnover rate in these early-stage manufacturing firms. The authors find that particularly when the constraints to self-employment were addressed through cash grants the workers preferred self-employment. 
      1. Does this mean that industrialization is not the best way to raise wages and lower the poverty rate in low-income countries? No. But it indicates that there may be a more efficient equilibria where a set of regulations providing a baseline level of safety for workers that address the issues identified in this study (chemical fumes, repetitive stress injuries, and probability of serious injury) can be beneficial to both employers via a lower turnover rate and to workers who would more likely work there if these health concerns were addressed. Such a set of regulations need not be so stringent that they reduce the comparative advantage of setting up shop in sub-Saharan Africa given the low wages on the continent but they will provide better standards of living for workers expected to drive these changes. 
Gender in the profession

On the panel on gender in the economics profession. The community by now is well aware of statistics indicating the low proportion of women who study economics as undergraduates, the lower proportion who study it as PhD candidates, and the even lower proportion who are tenured faculty at universities. The primary questions now, in my opinion, are (1) whether members of the community believe that these statistics are indicative of gender bias (as opposed to differences in ability or preference between the genders); and (2) whether members of the community believe that they can and should take action to address this bias, particularly when it is implicit and particularly where it requires the buy-in of economists who are neither part of the problem nor the solution.

Several of the questions posed in the panel revolve around these ideas. First is the need for data and evidence that is reflective of implicit bias to indicate to said economists that there is a problem at hand. Erin Hengel's paper on publication records of male and female economists that I discussed last year and Alice Wu's paper on sexism within the Econ Job Market Rumors website which is informal but commonly used among academic economists for job postings and career advice (see this interview with Wu on this paper) are two examples of this type of evidence. This webpage put together by the UC Berkeley Women in Economics group offers other useful information.

From my own anecdotes and research experience within the Gender Innovation Lab at the World Bank, there are a few issues that I think are actionable to address:
  1. Role models and social networks among women 
  2. Gender gap in perceived abilities in STEM fields 
  3. Culture and implicit bias within the profession
Given that the third issue is probably the one that is most difficult to address I think it requires first the buy-in from the community that I mentioned above. Being aware of implicit bias and its effects on the community are important because they are needed to take the next steps. For example, one issue that was talked about in the panel is aggression in economics seminars. It likely impacts women more than men because women tend to do better in collaborative and non-aggressive environments and the aggression tends to be more often directed towards women than it does towards other men (e.g. see Wu's paper on EJMR). But suffice it to say, I think we would all do better - men and women alike - if we were all a bit kinder to one another without compromising the rigor of our work. Specifically, to both acknowledge that we can and should be able to communicate questions and criticisms without resorting to aggression and be willing to learn the techniques to do so. Same with being willing to learn the techniques to recognize and address implicit bias.  

I have been supported in my efforts by peers and role model figures - mostly male - that have been enthusiastic about my ability to succeed in this profession. I have been blessed in not only role models in professional and academic life but also partners in my personal life that have been the most influential factors in my decision to undertake graduate studies. My thoughts on this issue are - in addition to addressing systematic issues within the field - if you can support a young person and believe in their abilities it is probably a determining factor in their decision to pursue higher studies. Whether we have the data or not as of yet (and there is more empirical research being conducted on role model figures and mentoring), we can't underestimate the value of empathy in how people decide whether or not they want to be in a particular location, field, university, firm. 

Wednesday, October 3, 2018

In-depth look at income and wealth data (pt. 2 of 3): Wealth

First, to preface with why wealth as distinct from income is relevant to economists and to policymakers at large. Kopczuk (2014) discusses the importance of understanding the wealth distribution: "the extent to which the well-off are going to rely on work vs. return to their wealth in the future is clearly important for assessing the extent to which a society will view itself in some way a meritocracy." Wealth is an important determinant of labor force participation and therefore impacts productivity and economic growth. It also has important implications for inequality, intergenerational mobility, and, consequently, implications for democratic institutions whose stability is reliant on a meritocratic society or at least the verisimilitude of a meritocratic society.

It should be noted that estimates of wealth inequality and the top wealth shares are not as widely agreed upon as estimates of income inequality and labor income shares. There are a few main data sources for estimating wealth inequality that are aptly summarized in Alvaredo, Atkinson, and Morelli (2018):
  1. Household surveys including the U.K. Wealth and Assets Survey and the U.S. Survey of Consumer Finances;
  2. Administrative data on individual estates at death; 
  3. Administrative data on wealth of living from annual wealth taxes; 
  4. Administrative data on investment income that are capitalized; and 
  5. Lists of large wealth-holders (e.g. Forbes).
These data sources are discussed in great detail in Kopczuk (2014)'s "What Do We Know About the Evolution of Top Wealth Shares in the United States?" which specifically discusses the U.S. Survey of Consumer Finances (1), the mortality multiplier method with individual estate data (2), and investment income data (4). Each of these data sources is subject to different concerns. Household surveys and list of the wealthiest individuals are recent phenomena and cannot be used for estimates prior to the 1950s when the household surveys on wealth were first implemented. Administrative data on wealth of the living based on wealth taxes cannot be recouped in most developed countries because only a few developed countries, most notably France and Norway, have a wealth tax to begin with. Therefore, most researchers rely on estate tax records on individual estates at death or on reported taxable capital income.

The primary concern with estate taxes is that the distribution of estates of the deceased must be projected to the population at large: i.e. a multiplier method must be used in order to answer the question, how does the distribution of wealth among the deceased reflect the distribution of wealth among the living? Mortality multipliers are inverses of mortality rates based on various criteria, for example, wealthy individuals tend to have lower mortality rates and increased longevity compared to less wealthy individuals and therefore a higher mortality multiplier would be applied to the upper estate ranges meaning there are relatively more individuals living within those ranges than lower ones. For more on recent discussions of the relative longevity of the wealthy see Saez and Zucman (2016) and Chetty et al. (2016).

Kopczuk presents a few interesting stylized facts about wealth that provide a good introduction to the wealth distribution and methods of estimating it:
  • Wealth is highly concentrated (top 10 percent holds between 65 and 85 percent of the total wealth, top 1 percent holds between 20 and 45 percent of total wealth based on time period); 
  • While the methods of estimating the wealth distribution disagree on the timing it is clear that wealth concentration hit its apex prior to the Great Depression and declined after that; 
  • Different methods lead to varying estimates for the top 1% for several reasons: one is that the estate tax multiplier method uses the individual as the unit of observation, surveys use the household, and the capitalization method uses tax units; another is that tax evasion impacts the administrative tax-based methods (estate tax and capitalization) but not the survey-based methods. Some capture debt (estate tax returns) whereas others do not (capitalization). 
In a recent issue of the Journal of Public Economics commemorating Tony Atkinson's work, Alvaredo, Atkinson, and Morelli (2018) provides new evidence on the evolution of top wealth shares in the U.K. To choose one of the most interesting facets of the discussion of wealth that they present in the article, it is enlightening to view the top wealth shares compared to the wealth shares excluding housing.



The top 1%'s total wealth share and wealth share excluding housing tracked each other for much of the late 20th century but the authors note the divergence between the two trends in the 21st century, wherein the share of the top 1% of wealth holders of total wealth increased much more rapidly than its share of wealth excluding housing. In other words, the growth of wealth excluding housing is likely to be a more significant contributor to rising inequality than is the growth of housing wealth. In fact, they even mention that increases in housing prices serve an equalizing effect for the top 1%:

"It appears that housing wealth has moderated a definite tendency for there to be a rise in recent years in top shares in total wealth apart from housing. When people talk about rising wealth concentration in the U.K., then it is probably the latter that they have in mind... The results show how the impact of a general rise in house prices has changed over the period but it is always equalizing for the top 1%. At the beginning of the period a rise of 25% led to a reduction of some 1 percentage point in the share of the top 1% but the effect became smaller over time."

It should be noted, however, that trends in the housing market - particularly the resurgence in the private landlord and "buy to let" over the past three decades - likely have impacts on other areas of the wealth distribution apart from the top 1% of wealth owners (though these impacts are not addressed in this paper). This New York Times article from last year, for example, is a news feature that discusses the role that homeownership plays in propagating existing wealth and income inequalities. These topics and the lower rungs of the wealth distribution more broadly are areas for further investigation, but for the time being Alvaredo, Atkinson, and Morelli (2018) highlight how granularity in wealth data can be used to better identify the causes of growing wealth inequality over the past few decades and, while they utilize estate data and the mortality multiplier method in their analysis, can also be triangulated with other methods and data sources to form a more comprehensive understanding of the wealth distribution.

Sources
  1. Alvaredo, F., Atkinson, A., Morelli, S. (2018). Top wealth shares in the UK over more than a century. Journal of Public Economics.
  2. Kopczuk, W. (2014). What do we know about the evolution of top wealth shares in the United States? NBER Working Paper 20734.
  3. Chetty, R., Stepner, M., Abraham, S., Lin, S., Scuderi, B., Turner, N., Bergeron, A., Cutler, D. (2016) The association between income and life expectancy in the United States 2001-2014. Journal of American Medical Association.
  4. Saez, E., Zucman, G. (2016) The distribution of US wealth, capital income, and returns since 1913. Quarterly Journal of Economics

Tuesday, August 14, 2018

In-depth look at income and wealth data (pt. 1.5 of 3): Non-traditional data and machine learning approaches

While this was originally meant to be a three-part series on income and wealth data, it would have been an oversight to not include some discussion of the non-traditional data and machine learning approaches to collecting information on poverty. These data are particularly relevant in developing countries where traditional sources of data - administrative data and survey data - are not collected as widely, regularly, or thoroughly. This can be for several reasons: nationally representative surveys are expensive and the costs of data collection too high, challenges associated with data collection in conflict-affected areas (discussed in greater detail in a previous publication I worked on), and large proportions of the population are employed in the informal economy meaning there is little by way of administrative tax records at the lower end of the income distribution.

Yet, information on poverty is still needed in these countries to inform evidence-based policymaking by governments, international organizations, and non-profits. A brief article by researcher Joshua Blumenstock published a few years ago in Science, "Fighting Poverty with Data", discusses the frontier of research in this area that aims to supplement the traditional sources of data on wealth and inequality with machine learning approaches. Blumenstock discusses, for example, the rise in use of nightlight data to track economic productivity and growth citing one paper which utilizes nightlight based measures to study the impact of sanctions on North Korea. In fact, a paper that I reviewed earlier in the year on the impact of Chinese aid projects on local corruption used nightlight data to proxy for local economic activity in areas around active and inactive Chinese aid sites.

More novel and more interestingly, the author cites research in machine learning that uses satellite imagery in conjunction with nightlight data to identify the visual features of relatively wealthier areas (which have brighter nightlight) that would allow researchers to leverage daytime satellite images to better track poverty in developing countries. There are limitations to this approach for example that nightlight is not an ideal measure of activity at the lower end of the income distribution - where all is dark - but with further research these approaches could be very useful in the developing country context.

Mobile phone data - which was discussed in part in the above article and in greater detail in this other Science piece also by Blumenstock - is also promising. Using mobile phone logs, researchers extract statistics including volume, intensity, and timing of phone calls, the structure of the individual's network of contacts, and mobility and migration information based on geospatial markers and whittle down to the statistics that can be used to predict socioeconomic status. In the case cited in this article, the researchers paired consenting individuals' mobile phone data with survey data that they collected on individual income and wealth in order to train the model. It should be noted that mobile phone data is subject to greater ethical and privacy concerns than publicly available data. While the research cited here aimed to obtain macro level statistics to inform policymaking it is clear that attempting to obtain a more granular understanding for specific demographics will be challenging. ICT access and use is far from universal and, often, those who are excluded from its access are the most vulnerable. This is similar to the challenges with using conflict data wherein the data on those who are the most vulnerable and impacted by conflict is the data that is the most challenging to collect and to collect accurately. This is not, however, meant to generalize, given that some of the poorest regions of the world have reasonably high mobile phone penetration but rather a cautionary note when assessing whether data are representative with respect to specific populations.

For example, with respect to a recent project that I've worked on, there is high mobile phone penetration in sub-Saharan Africa despite low income. Yet, while its neighbors in East Africa have experienced fast growth in mobile phone ownership and usage in the past five years, Ethiopia has fallen behind largely due to government ownership of the nation's telecom monopoly which has limited expansion and service. Further, analyzing the distributional data on mobile phone usage indicates that women are far less likely to own and use mobile phones than men - consistent with the findings in many developing countries - and that any data collected from these devices in a hypothetical scenario would only be representative of a specific demographic.

And yet, despite the challenges, non-traditional sources of data offer promise particularly in geographic areas where recent, traditional data on wealth and poverty are not available. Research in this interdisciplinary area will be interesting to watch in the near future.

Tuesday, July 3, 2018

In-depth look at income and wealth data (pt. 1 of 3): Background

For some time now I have been interested in writing an in-depth post on income and wealth data in order to discuss how the study of inequality - in conjunction with the data and methods that enable this study - has progressed over time. While this was initially intended to be a single post, it quickly became evident that there was too much to discuss within too short a space. In this first post of a three-part series, then, I focus on providing the background for a more granular discussion of wealth and income in the next two parts.

Given the topic at hand, it is noteworthy that several articles on inequality and tax and redistribution policy were published in a recent special issue of the Journal of Public Economics honoring the late Tony Atkinson. For an introduction to that series of papers see here. My previous post on individual and household level inequality is based on a paper within this special issue. Additionally, a recent issue of the Quarterly Journal of Economics features an article that combines national accounts data with micro data to produce estimates of inequality in the U.S. that are consistent at the macro level.

In light of expanding research on inequality, its growing presence in policy debates in developed countries, and the evolution of both data and methods that enable its rigorous study, it is useful to take stock of the existing data sets and methods used by researchers to answer some of the most pressing questions in public economics today: those that deal with the distribution of wealth and income in our societies and the reasons for widening or stagnant inequality levels. We can also assess what types of questions we are now able to answer and how our answers to these and other - yet unasked - questions can become more accurate through improved data collection and methods and how future data collection can fill existing gaps in our knowledge.

To begin, the World Inequality Database - a database of global wealth and income inequality data co-founded by Tony Atkinson - provides a concise description of data and research in this field over the past twenty years. Two important trends:
  1. Most studies on inequality have until very recently focused on income rather than wealth. The key reason is the greater availability of micro data to study income, which is taxed and therefore observable in administrative data, as opposed to wealth, which in most developed countries is not taxed apart from an estate tax upon death. A secondary reason is that it has not been made evident until recently - likely for similar data reasons - that wealth concentration plays a large role in the inequality we see within developed countries. Piketty (2014)'s Capital in the Twenty-First Century was not the first but perhaps the most prominent description of the growing role of capital in widening divisions between haves and have-nots.
  2. Current efforts are aimed at producing distributed national accounts that combine administrative micro data with national accounts macro data - ledgers of assets and liabilities at the national level - in order to reconcile inequality estimates that are created based on micro data with the national accounting. This publication from the founders of the WID discusses the motivation and methodology for the creation of these "distributed national accounts." It notes the historical background, "[by] combining the macro and micro dimensions of economic measurement, we are of course following a very long tradition. In particular, it is worth recalling that Kuznets was both of the founders of the U.S. national accounts and the author of the first national income series and also the first scholar to combine national income series and income tax data in order to estimate the evolution of the share of total income going to top fractiles in the U.S. over the 1913-1948 period (see Kuznets, 1953)." The article cited above from the QJE, Piketty, Saez, and Zucman (2018), presents "distributed national accounts" for the U.S., which they note is distinct from government statistical agencies' work in this area.
Discussion of the main data types and their roles in inequality studies

Administrative micro data

To preface a discussion on administrative tax data for wealth and income studies, I provide context for use of this data for social sciences research more broadly. Administrative data are collected for the purposes of registration, transaction and record keeping, and are often linked to public service delivery. They are typically collected by public sector agencies and can be used in administration systems in education, health, and taxation, among other departments of the public sector. It should be noted that these data are "found" data and are not collected for the purposes of research as survey data are. The social sciences, and economics in particular, have shifted to using administrative data over survey data sources in recent years for several reasons.

Specifically as noted in this white paper to the National Science Foundation: "Administrative data are highly preferable to survey data along three key dimensions. First, since full population files are generally available, administrative records offer much larger sample sizes... Second, administrative files have an inherent longitudinal structure that enables researchers to follow individuals over time and address many critical policy questions, such as the long term effects of job loss (von Wachter, Song, and Manchester, 2009) or the degree of earnings mobility over the life cycle (Kopczuk, Saez, and Song, 2010). Third administrative data provide much higher quality information than is typically available for survey sources, which suffer from high and rising rates of non-response, attrition, and under-reporting."

Access to this data is not without its challenges in many developed countries. Nordic countries have been leaders in enabling researchers to access de-identified administrative or "register" data but other countries, such as the U.S., have been relatively slow to follow. Given the central role that administrative data has come to play in social sciences and economics research in particular (see the two charts on the number of publications in leading economics journals that employed administrative data in this presentation from researcher Raj Chetty, who also co-authored the white paper cited above), it is clear that access to these data has important implications that are outlined in an article published in the Economist last month on the topic.

Administrative tax data are widely used in income and wealth inequality studies. For example, wealth inequality is largely studied through either estate tax records - in order to create wealth distributions of wealth at death and to extrapolate from those records the distribution of wealth among the living using the mortality multiplier method - or through taxable capital income (it should be noted that only one-third of total capital income is reported on tax returns which is why it is challenging to estimate wealth based on this quantity). Similarly, income inequality is studied through income tax records. Given the socioeconomic and demographic data contained in these records we are able to answer (or attempt to answer) a wide range of social science research questions based on micro data. Yet, the missing piece is information on movements in the economy at large over time (e.g. increase in fraction of retired individuals or declines in household size) which could have implications for inequality.

As noted by Piketty, Saez, and Zucman (2018), studies that use micro data exclusively are unable to answer questions such as: (1) what fraction of economic growth accrues to different parts of the income distribution, (2) what fraction of the increase in income inequality is due to changes in share of labor vs. capital in national income as opposed to changes in the distribution within labor or capital earnings, (3) how does government redistribution impact inequality (i.e. we are only able to observe pretax income using micro data series which does not allow us to observe the changes in the income distribution between pre- and posttax). To answer these questions, they argue, merging micro data with national accounts data at the macro level is valuable.

National accounts macro data

On the macro side side, national accounts data aggregate output, expenditure, and income activities of each sector of the economy. While income and consumption measures are important for evaluating standards of living they offer only a static picture of well-being. Specifically, income and consumption reflect current well-being: how much a household or an economy is producing and consuming at present, but they do not provide much insight into a household or economy's long-term or future well-being (beyond making assumptions that current well-being and consumption are highly correlated over time). This is where national accounts data can be useful: data on a household or economy's ownership of marketable assets and contraction of debts can provide insight into long-term or future well-being though it may be cross-sectional rather than longitudinal.

For a valuable introduction to balance sheets and the national accounts data see here for a discussion from the French National Institute of Statistics and Economic Studies (INSEE). It should be noted that the definitions of "assets" and whether or not they provide "economic advantages" refer specifically to those items that have market values. This would exclude, as stated by INSEE, "items that one might expect to see in the accounts (human capital, natural heritage, natural State property, household durables, pension entitlements linked to the allocation system, etc.)" They note as a rule of thumb that only items that are featured in the capital and financial accounts are included as assets in order to maintain internal consistency. The capital account and financial account link the opening and closing balance sheets to one another: they specify what happened to the accumulation of capital based on capital consumption, assets sold and acquired, discoveries and inventions, and nominal holding gains as a result of price fluctuations.

These data, and specifically the national income measures in these data, may be relied upon to fill the gaps in our knowledge from the tax data. Specifically, there are gaps between the reported income and the national income that are not captured in micro studies: imputed rents of homeowners and taxes on top of unreported and untaxed labor income in the form of tax-exempt fringe benefit. Piketty, Saez, and Zucman (2018) estimate that the fraction of national income reported on tax returns in the U.S. has declined from 70 percent in the late 1970s to roughly 60 percent today which indicates that micro data alone may underestimate the level and growth of income in this country and perhaps more so for certain parts of the income distribution than others depending on what exactly is being excluded from the tax data that is present in the national income.

For a more in-depth description of the methods and the process by which these two data are being combined, I would look to the article. The authors effectively illustrate both the motivation and the methods for incorporating national income macro data into inequality studies. In the next part of this three-part series, I will discuss the data and research on wealth inequality specifically to provide greater detail on wealth estimates using estate tax data compared to those using capital income.

Sources
  1. Piketty, T., Saez, E., Zucman, G. (2018). Distributional National Accounts: Methods and Estimates for the United States. Quarterly Journal of Economics.
  2. Kleven, H., Luttmer, E. (2018). A Special Issue of the Journal of Public Economics: Honoring the Work of Sir Anthony B. Atkinson (1944-2017). Journal of Public Economics. 
  3. Blundell, R., Joyce, R., Keiller, A.N., Ziliak, J.P. (2017). Income inequality and the labour market in Britain and the US. Journal of Public Economics.