Brett Keller, MPA
Back in September the New York Times reported on an unexpected finding from a clinical trial: “A promising but expensive device to prop open blocked arteries in the brain in the hope of preventing disabling or fatal strokes failed in a rigorous study.” Many promising medical innovations fall short when they finally reach clinical trials, but this story was unusual because the stents had already been approved by the FDA under a so-called humanitarian exemption. The FDA approved the stents to reduce the risk of stroke, but those who received it had twice as many strokes.
How did this happen? The Times chronicled experts’ puzzlement: “Researchers said the device seemed as if it should work.” And Joseph Broderick, a prominent neurologist, is quoted as saying “Quite frankly, the results were a surprise.” Researchers are delving into this case to discover why the stent failed, but policymakers from all fields should take it as a valuable lesson. This is one more argument for testing policies whenever possible: not only does expert opinion sometimes get things wrong, but without good data there is often no way to really know when they are right.
Similar lessons can be gleaned from the history of surgical response to breast cancer. In The Emperor of All Maladies (2010), a new history of cancer, oncologist Siddhartha Mukherjee chronicles the history of such failed interventions as the radical mastectomy. Over a period of decades this brutal procedure – removing the breasts, lymph nodes, and much of the chest muscles – became the tool of choice for surgeons treating breast cancer. In the 1970s rigorous trials comparing radical mastectomy to more limited procedures showed that this terribly disfiguring procedure did not in fact help patients live longer at all. Some surgeons refused to believe the evidence – to believe it would have required them to acknowledge the harm they had done. But eventually the radical mastectomy fell from favor; today it is quite rare. Many similar stories are included in a free e-book titled Testing Treatments (2011).
As a society we’ve come to accept that medical devices should be tested by the most rigorous and neutral means possible, because the stakes are life and death for all of us. Thousands of people faced with deadly illnesses volunteer for clinical trials every year. Some of them survive while others do not, but as a society we are better off when we know what actually works. For every downside, like the delay of a promising treatment until evidence is gathered properly, there is an upside – something we otherwise would have thought is a good idea is revealed not to be helpful at all.
Under normal circumstances most new drugs are weeded out as they face a gauntlet of tests for safety and efficacy required before FDA licensure. The stories of the humanitarian-exemption stent and the radical mastectomy are different because these procedures became more widely used before there was rigorous evidence that they helped, though in both cases there were plenty of anecdotes, case studies, and small or non-controlled studies that made it look like they did. This haphazard, post-hoc testing is analogous to how policy in many other fields, from welfare and education, is developed. Many public policy decisions have considerable impacts on our livelihoods, education, and health. Why are we note similarly outraged by poor standards of evidence that leads to poor outcomes in other fields?
A recent example from New York City helps illustrate how helpful good evidence can be in shaping policy. A few years ago Mayor Michael Bloomberg rolled out a massive program that seemed to make a lot of sense: pay teachers bonuses based on their students’ performance. The common sense proposal was hailed as “transcendent” and gained the support of the teachers’ union. It cost $75 million, and it didn’t work. How do we know? The program was designed from the beginning as a pilot where schools were randomly assigned to the program or to a control group, and the research showing that the program had no effect on outcomes was subsequently published. What would have happened if this policy had been put in place without an effective evaluation plan? In all likelihood New York officials would now be touting its success at conferences and urging other cites to implement similar programs. Instead it was quietly shelved. That this particular program did not have the intended effect is disappointing, but it is much better than if we believed it worked and continued on unaware.
The pros and cons of randomized trials have been discussed here on 14 Points before – see recent posts by Jake Velker and Shawn Powers. The cases I presented here are ones where the results were not “no-brainers” at all, and without systematic evaluation bad policies would have been or tragically were put in place. While good evidence does not have to come from randomized trials, there are still many areas where they are underused. In areas where they are feasible (i.e. not macroeconomics) such evidence should be the norm, and those who implement policies with great optimism but without planning for thoughtful evaluation should be panned. Even without random assignment of the treatment, the best policy evaluations should involve a serious attempt to estimate the counterfactual: what would have happened in the absence of the intervention. Moving beyond arguments over specific programs and whether they work, policymakers can move us towards better outcomes by creating a culture where strong evidence is valued. After all, the clinical trial as we know it in medicine is a 20th century innovation; it hasn’t always been this way.
A student-run public policy blog of the Woodrow Wilson School of Public and International Affairs at Princeton University.
NOTE: The views expressed here belong to the individual contributors and not to Princeton University or the Woodrow Wilson School of Public and International Affairs.
Showing posts with label methods. Show all posts
Showing posts with label methods. Show all posts
Friday, December 16, 2011
Friday, November 25, 2011
Empowering Evaluation: Looking beyond the numbers
Elizabeth Hoody, MPA
As a former grant-maker, I was glad to see that 14 Points talked about the challenge of evaluating anti-domestic violence work in a post by Payal Hathi last May. In my own experiences working in a women’s rights foundation, I saw just how difficult it can be to quantify the impact of women’s rights organizing in numerical terms. While numbers can tell an important story (such as how many young women receive sexual and reproductive health education), they often leave out what for me is most compelling about a group’s work. This might be the reflections of an individual young woman who now feels that she can talk to her partner about contraceptives or the story of a group of girls who decided to form their own anti-trafficking student organization after participating in a prevention workshop. So while there are many valid and pressing questions about how to “get the numbers right” in program evaluation, my bigger concern these days is how to evaluate impact beyond the numbers…and then again how to aggregate and share this type of evaluation in a way that donors, policymakers, and peer organizations can easily understand.
As a former grant-maker, I was glad to see that 14 Points talked about the challenge of evaluating anti-domestic violence work in a post by Payal Hathi last May. In my own experiences working in a women’s rights foundation, I saw just how difficult it can be to quantify the impact of women’s rights organizing in numerical terms. While numbers can tell an important story (such as how many young women receive sexual and reproductive health education), they often leave out what for me is most compelling about a group’s work. This might be the reflections of an individual young woman who now feels that she can talk to her partner about contraceptives or the story of a group of girls who decided to form their own anti-trafficking student organization after participating in a prevention workshop. So while there are many valid and pressing questions about how to “get the numbers right” in program evaluation, my bigger concern these days is how to evaluate impact beyond the numbers…and then again how to aggregate and share this type of evaluation in a way that donors, policymakers, and peer organizations can easily understand.
In the past year, I have come across several creative examples of evaluation strategies that do attempt to move beyond the numbers. Many of these strategies attempt to articulate, verbally or visually, the systemic impact of an organization’s work. One example is a map that was released by the Global Fund for Women just this past week, which captures the Fund’s impact around the world using bright spots. In the Global Fund’s words, this map “explores where a relationship between Global Fund for Women and grantee groups is more likely to yield a higher movement building impact.” While the map does rely on a series of numerical indicators, the visual analysis tells a bigger story about the collective impact of Global Fund for Women grants on women’s rights movements around the world.
The second example is the Gender At Work Framework, which “helps organizations see their work from new perspectives by combining best practices in organizational development with feminist thought.” One tool that the framework uses is a graph where civil society organizations can plot the different types of social changes they are addressing through their work. The graph places the continuum of “individual versus systemic” change on vertical axis and “formal vs. informal” changes on the horizontal axis, resulting in four quadrants of change:
- Women’s access to resources (quadrant I)
- Women’s and men’s consciousness (quadrant II)
- Informal cultural norms and exclusionary practices (quadrant III)
- Formal institutions, laws, and policies (quadrant IV).
Women’s organizations use the tool to visually represent the changes that they are trying to impact through their work. The graph is also a good indicator of what programs will be easier to quantitatively evaluate (such as programs that work for specific policy changes) and those that will be difficult to track through traditional metrics (how do you measure changes in men’s and women’s consciousness?).
What I like most about the Gender At Work Framework is that it encourages grassroots organizations to define evaluation framework themselves at the beginning of their planning processes. When this happens, evaluation shifts from being a chore for donors to being an effective way of tracking and reflecting on an organization’s progress towards its goals. In my personal, unquantifiable opinion, the resulting story more often than not contains richer analysis, more honest reflections, and genuine learning.
For the past year, I have been working on the Advisory Committee of FRIDA – The Young Feminist Fund, which is a newly formed foundation that is run for and by young feminist activists. As FRIDA prepares for its first grant-making round, these questions about evaluation are on my mind. How do we tell our story to our first funders and how do we empower the first FRIDA grantees to take ownership over evaluation of their work? Can we be accountable and collaborative in evaluation? As a starting point, we are planning to develop the first grant evaluation process in partnership with FRIDA’s first round of grantees. As we begin to hash out exactly what the process will look like, I am grateful for these new tools that push us towards more creative, dynamic ways of articulating our goals, visions, and impacts.
Friday, November 11, 2011
Vying with Velker: RCTs reconsidered
Shawn Powers, MPA ’11
In “Randomized Controlled Trials on Trial,” Jake Velker proposes several reasons to be skeptical of randomized controlled trials (RCTs) as a method of program evaluation. While Jake makes some good points, as a “randomista” I think the picture he paints of the RCT movement is far too pessimistic. I will consider each of his arguments in turn.
1. External validity
Jake first takes up the critique, championed by Princeton’s Angus Deaton, that RCTs suffer from external validity problems—in other words, the results of an evaluation may not generalize well to other contexts. While the criticism is frequently leveled at RCTs, the question of external validity applies to all empirical work. In general, I find these discussions about the differences between Deaton and RCT proponents a bit overblown. Perhaps this is because everyone loves a good spat between high-profile intellectuals (see also: Sachs and Easterly). In reality, according to the profile of Esther Duflo in The New Yorker that Jake references, Deaton has described his criticisms “as more in the form of an amicus brief than an attack." The arguments raised by Deaton and others are having an impact, as Jake acknowledges, and RCTs increasingly are testing rich behavioral hypotheses. The incentives of academic publication are also moving RCTs toward greater theoretical sophistication. Gone are the days when a randomized design was a novel enough identification strategy that it could propel a study to publication in a top economics journal.
Considerations of theory aside, whether or not a particular result generalizes is itself a testable, empirical question. We cannot—and should not—test everything everywhere, but if a particular approach proves effective in multiple contexts, our confidence in its “generalizability” should increase accordingly. Replication studies can also test variations in the length or intensity of treatment, disentangle the impact of different components of a program, or test how a small-scale intervention performs as it is scaled up.
If the goal is to achieve certainty that intervention X will achieve result Y in context Z, we will never achieve it, with RCTs or any other method. However, considering evidence from even one rigorous evaluation is a big improvement over flying blind. As we consider multiple evaluations, together with insights from theory and other empirical work, the picture becomes that much clearer.
2. Institutional constraints
Jake’s main criticism is that economists conducting RCTs “have been accused of ignoring the institutional constraints against which their interventions would inevitably contend if scaled up.” He cites corrupt bureaucracies, weak institutions, a lack of (or perverse) performance incentives, and budgetary problems as barriers to successful replication of programs found to be effective. The underlying message seems to be that if the RCT movement wants to influence policy successfully, it cannot just publish research findings and hope for the best.
I could not agree more with this last point, but the RCT community is much farther along on this front than Jake suggests. Both J-PAL and our sister organization, Innovations for Poverty Action (IPA), have policy staff dedicated to bringing research findings to bear on the often-messy world of policymaking. While our academic affiliates are involved with policy outreach, non-academic policy staff help extend the reach of their research findings. This process is never easy and not always successful, for all the reasons Jake mentions, but we have found that it is possible to improve policy even in very constrained environments.
As an aside, Jake also suggests that governance in developing countries is not amenable to quantitative study. I would have thought the same before I started with J-PAL, but in fact, J-PAL affiliates currently have at least 42 completed or ongoing evaluations in political economy and governance, including many that address precisely the issue he raises of the incentives of government officials and service providers.
3. Do we already know what works?
Finally, Jake entertains the idea that perhaps we already know what works, since “many of the most celebrated finds of the RCT movement are relative ‘no-brainers’.” I find this argument troubling for two reasons. First, we have been following our intuition about what works in development for decades, with not a lot to show for it. What we have seen is succession of fads, with decidedly mixed results in terms of reducing poverty. There was a time when infrastructure was the “no brainer,” later it was basic needs, still later the focus turned to sustainable development, and today infrastructure seems back in vogue. To suggest that we already know what to do invites just this kind of intellectual drift.
Second, while it may be true that many RCTs report seemingly obvious findings, some of them surprise us—and we never know in advance which those will be. For example, a number of NGOs and opinion leaders have championed the idea that distributing sanitary products to adolescent girls will remove a barrier to female education. The underlying common-sense assumption is that menstruation causes many missed days of school. However, a randomized evaluation of a program that distributed an easy-to-use sanitary product in Nepal found no significant effect on school attendance (although the girls used, and liked, the product). As always, we should avoid over-generalizing from one study, but at minimum these findings suggest that proponents of this approach should adjust their expectations about what it can deliver. In other cases, RCTs have contributed clear evidence to debates where both camps have common-sense arguments on their side, such as the vexed issue of whether, and how much, to charge poor people for basic health and education products and services. Finally, even if the qualitative findings of an RCT seem to confirm common sense, policymakers may still want to know how an intervention stacks up quantitatively against other interventions with the same goal, in terms of both raw impact and cost-effectiveness.
RCTs are no more a panacea for development than anything that came before them. But as long as there is more ideology and wishful thinking in development policymaking than evidence, I believe that the continuing growth of the RCT movement is a welcome trend.
Shawn Powers is a Policy Manager at the Abdul Latif Jameel Poverty Action Lab (J-PAL). The opinions expressed here are his own.
In “Randomized Controlled Trials on Trial,” Jake Velker proposes several reasons to be skeptical of randomized controlled trials (RCTs) as a method of program evaluation. While Jake makes some good points, as a “randomista” I think the picture he paints of the RCT movement is far too pessimistic. I will consider each of his arguments in turn.
1. External validity
Jake first takes up the critique, championed by Princeton’s Angus Deaton, that RCTs suffer from external validity problems—in other words, the results of an evaluation may not generalize well to other contexts. While the criticism is frequently leveled at RCTs, the question of external validity applies to all empirical work. In general, I find these discussions about the differences between Deaton and RCT proponents a bit overblown. Perhaps this is because everyone loves a good spat between high-profile intellectuals (see also: Sachs and Easterly). In reality, according to the profile of Esther Duflo in The New Yorker that Jake references, Deaton has described his criticisms “as more in the form of an amicus brief than an attack." The arguments raised by Deaton and others are having an impact, as Jake acknowledges, and RCTs increasingly are testing rich behavioral hypotheses. The incentives of academic publication are also moving RCTs toward greater theoretical sophistication. Gone are the days when a randomized design was a novel enough identification strategy that it could propel a study to publication in a top economics journal.
Considerations of theory aside, whether or not a particular result generalizes is itself a testable, empirical question. We cannot—and should not—test everything everywhere, but if a particular approach proves effective in multiple contexts, our confidence in its “generalizability” should increase accordingly. Replication studies can also test variations in the length or intensity of treatment, disentangle the impact of different components of a program, or test how a small-scale intervention performs as it is scaled up.
If the goal is to achieve certainty that intervention X will achieve result Y in context Z, we will never achieve it, with RCTs or any other method. However, considering evidence from even one rigorous evaluation is a big improvement over flying blind. As we consider multiple evaluations, together with insights from theory and other empirical work, the picture becomes that much clearer.
2. Institutional constraints
Jake’s main criticism is that economists conducting RCTs “have been accused of ignoring the institutional constraints against which their interventions would inevitably contend if scaled up.” He cites corrupt bureaucracies, weak institutions, a lack of (or perverse) performance incentives, and budgetary problems as barriers to successful replication of programs found to be effective. The underlying message seems to be that if the RCT movement wants to influence policy successfully, it cannot just publish research findings and hope for the best.
I could not agree more with this last point, but the RCT community is much farther along on this front than Jake suggests. Both J-PAL and our sister organization, Innovations for Poverty Action (IPA), have policy staff dedicated to bringing research findings to bear on the often-messy world of policymaking. While our academic affiliates are involved with policy outreach, non-academic policy staff help extend the reach of their research findings. This process is never easy and not always successful, for all the reasons Jake mentions, but we have found that it is possible to improve policy even in very constrained environments.
As an aside, Jake also suggests that governance in developing countries is not amenable to quantitative study. I would have thought the same before I started with J-PAL, but in fact, J-PAL affiliates currently have at least 42 completed or ongoing evaluations in political economy and governance, including many that address precisely the issue he raises of the incentives of government officials and service providers.
3. Do we already know what works?
Finally, Jake entertains the idea that perhaps we already know what works, since “many of the most celebrated finds of the RCT movement are relative ‘no-brainers’.” I find this argument troubling for two reasons. First, we have been following our intuition about what works in development for decades, with not a lot to show for it. What we have seen is succession of fads, with decidedly mixed results in terms of reducing poverty. There was a time when infrastructure was the “no brainer,” later it was basic needs, still later the focus turned to sustainable development, and today infrastructure seems back in vogue. To suggest that we already know what to do invites just this kind of intellectual drift.
Second, while it may be true that many RCTs report seemingly obvious findings, some of them surprise us—and we never know in advance which those will be. For example, a number of NGOs and opinion leaders have championed the idea that distributing sanitary products to adolescent girls will remove a barrier to female education. The underlying common-sense assumption is that menstruation causes many missed days of school. However, a randomized evaluation of a program that distributed an easy-to-use sanitary product in Nepal found no significant effect on school attendance (although the girls used, and liked, the product). As always, we should avoid over-generalizing from one study, but at minimum these findings suggest that proponents of this approach should adjust their expectations about what it can deliver. In other cases, RCTs have contributed clear evidence to debates where both camps have common-sense arguments on their side, such as the vexed issue of whether, and how much, to charge poor people for basic health and education products and services. Finally, even if the qualitative findings of an RCT seem to confirm common sense, policymakers may still want to know how an intervention stacks up quantitatively against other interventions with the same goal, in terms of both raw impact and cost-effectiveness.
RCTs are no more a panacea for development than anything that came before them. But as long as there is more ideology and wishful thinking in development policymaking than evidence, I believe that the continuing growth of the RCT movement is a welcome trend.
Shawn Powers is a Policy Manager at the Abdul Latif Jameel Poverty Action Lab (J-PAL). The opinions expressed here are his own.
Thursday, October 27, 2011
What Do We Do With Data Soup?
Katherine DiSalvo, MPA
As policy professionals, we’re likely to encounter messes of contradictory findings more and more throughout our careers.
There is contradictory data on important issues like the real level of US poverty, whether moving people out of a neighborhood of concentrated poverty improves their chances in life, the success of charter schools, or the effectiveness of giving away free bed-nets to combat malaria.
Do you know what to do with data soup? At the Woodrow Wilson School, I don’t think we students learn this sufficiently.
According to R. Kent Weaver, in Ending Welfare as We Know It (Brookings, 2000), the 1980s and 1990s saw a “multiplication” of policy research with “differing assumptions and conclusions.” Simultaneously, interest groups were adopting social science techniques and creating “a welter of conflicting findings.” In a separate article Weaver and a colleague assert that this may result in the “devaluation of the currency” of policy research. Weaver argues that it may “cause legislators to simply dismiss all evidence that does not fit their personal or constituency preferences.”
Devaluation of policy research is becoming commonplace. Even in the era of “data-driven” education leadership, Brenda Welburn, the head of the National Association of State Boards of Education (NASBE), recently told a WWS workshop team researching “School Choice and Impacts on Cities” that State Boards of Education members don’t know whose data to trust. As board members attempt to make state education policy and funding decisions, sometimes they don’t know how to do it with facts. “We [at NASBE] are dealing with perceptions, often,” Welburn said.
I don’t think it’s easy to digest data soup, and I think the Woodrow Wilson School needs to do more to help its students develop this ability. You may scoff and tell me you know how to wade through the stew. You know statistics! You know what research methods matter!
I don’t think any policy professional can rely on statistical prowess alone. The statistics program at the Wilson School is strong, and its decision to expand statistics requirements was a good one. However, with our limited time we students (not to mention professionals) cannot dig into data sets, look at assumptions, and evaluate every conclusion we read for ourselves. While some such analysis might be possible before an important policy decision or publication, we consume too much information to scrutinize it all.
The best proof that policy students won’t always use technical skills to sort through conflicting data professionally is that Woodrow Wilson students don’t always do so here! When I encounter conflicting data in classes, I’m too often told we students should dig deeper and decide who’s right…later.
We can’t rely exclusively on the “the gold standard” professors teach us to love: data generated by randomized control trials (RCTs). This creates an easy top tier of information on too few topics. Additionally, all the emphasis we hear on the “gold standard” may lead us to trust in RCT-based research too easily. The best part of the WWS course on data-based decision making is hearing Professor Lorenzo Moreno talk about how complicated it can be to do the right thing in the evaluation field. All that glitters…
We policy students need more practice criticizing questionable research. We need more practice wading through data mess and taking and defending a stand – not on politics, as we do in the introductory 501 course, Politics and Public Policy, but a stand on what we think is the truth. We need more sophisticated conversations about what data to trust and about how to evaluate vendors of policy research when we cannot evaluate each product. We need more shorthand than one “gold” standard.
We also need to talk about making policy in a world where different “facts” are consumed by different constituencies, and the truth is always up for debate. It’s the world in which we live, and it’s likely to get worse. If the Woodrow Wilson School could prepare us to digest data soup and to help change these cooking trends, that would truly be in the nation’s service and in the service of all nations.
As policy professionals, we’re likely to encounter messes of contradictory findings more and more throughout our careers.
There is contradictory data on important issues like the real level of US poverty, whether moving people out of a neighborhood of concentrated poverty improves their chances in life, the success of charter schools, or the effectiveness of giving away free bed-nets to combat malaria.
Do you know what to do with data soup? At the Woodrow Wilson School, I don’t think we students learn this sufficiently.
According to R. Kent Weaver, in Ending Welfare as We Know It (Brookings, 2000), the 1980s and 1990s saw a “multiplication” of policy research with “differing assumptions and conclusions.” Simultaneously, interest groups were adopting social science techniques and creating “a welter of conflicting findings.” In a separate article Weaver and a colleague assert that this may result in the “devaluation of the currency” of policy research. Weaver argues that it may “cause legislators to simply dismiss all evidence that does not fit their personal or constituency preferences.”
Devaluation of policy research is becoming commonplace. Even in the era of “data-driven” education leadership, Brenda Welburn, the head of the National Association of State Boards of Education (NASBE), recently told a WWS workshop team researching “School Choice and Impacts on Cities” that State Boards of Education members don’t know whose data to trust. As board members attempt to make state education policy and funding decisions, sometimes they don’t know how to do it with facts. “We [at NASBE] are dealing with perceptions, often,” Welburn said.
I don’t think it’s easy to digest data soup, and I think the Woodrow Wilson School needs to do more to help its students develop this ability. You may scoff and tell me you know how to wade through the stew. You know statistics! You know what research methods matter!
I don’t think any policy professional can rely on statistical prowess alone. The statistics program at the Wilson School is strong, and its decision to expand statistics requirements was a good one. However, with our limited time we students (not to mention professionals) cannot dig into data sets, look at assumptions, and evaluate every conclusion we read for ourselves. While some such analysis might be possible before an important policy decision or publication, we consume too much information to scrutinize it all.
The best proof that policy students won’t always use technical skills to sort through conflicting data professionally is that Woodrow Wilson students don’t always do so here! When I encounter conflicting data in classes, I’m too often told we students should dig deeper and decide who’s right…later.
We can’t rely exclusively on the “the gold standard” professors teach us to love: data generated by randomized control trials (RCTs). This creates an easy top tier of information on too few topics. Additionally, all the emphasis we hear on the “gold standard” may lead us to trust in RCT-based research too easily. The best part of the WWS course on data-based decision making is hearing Professor Lorenzo Moreno talk about how complicated it can be to do the right thing in the evaluation field. All that glitters…
We policy students need more practice criticizing questionable research. We need more practice wading through data mess and taking and defending a stand – not on politics, as we do in the introductory 501 course, Politics and Public Policy, but a stand on what we think is the truth. We need more sophisticated conversations about what data to trust and about how to evaluate vendors of policy research when we cannot evaluate each product. We need more shorthand than one “gold” standard.
We also need to talk about making policy in a world where different “facts” are consumed by different constituencies, and the truth is always up for debate. It’s the world in which we live, and it’s likely to get worse. If the Woodrow Wilson School could prepare us to digest data soup and to help change these cooking trends, that would truly be in the nation’s service and in the service of all nations.
Tags:
education,
Field IV (Economics),
methods
Sunday, October 2, 2011
Slum-free cities in India? The implementation realities of central government policies
Renee Ho, MPA
In 2009, as part of an effort to promote “slum-free cities,” the Indian Ministry of Housing and Urban Poverty Alleviation (HUPA) announced a new initiative, the Rajiv Awas Yojana (RAY).[1] The RAY aims to upgrade and bring all slums within the formal housing system. It also intends to prevent the development of new slums by providing more affordable housing options and developing more land to meet growing needs. As a prerequisite to RAY funding, states must first assign legal title to slum-dwellers over their living space.
In theory and on paper, the RAY includes a number of positive measures for slum-dwellers, especially the emphasis on property rights and the lack of distinction between government-recognized and unrecognized slums. If implemented properly, it could result in improved housing and access to services for many more of the urban poor.
However, in India there is a large disconnect between policy and implementation. The RAY Guidelines for Slum-free City Planning lack appropriate measures to ensure that slum-dwellers’ interests will be protected against arbitrary eviction and relocation, especially when private builders are involved in slum re-development. Without strict provisions for community participation, transparency, monitoring, and accountability, there is a danger that the RAY could be used to grab valuable slum lands in the center of the city while relocating poorer residents to faraway locations.
Potential problems with the RAY can be broken down into several categories:
1) Lack of Monitoring Mechanisms and Transparency
---------------
Notes
[1] While there is no exact translation of "Rajiv Awas Yojana," in Hindi "yojana" means "scheme," which in India is the word used for government programs or policies. "Awas" is commonly used for yojanas dealing with housing related issues. While Rajiv literally translates as "lotus flower," we think the term refers to something else. In India there is also a program called Indira Awas Yojana, started in 1985, to provide housing for the rural poor. Indira was the name of a female prime minister of India named Indira Gandhi, who served from 1966-1977 and from 1980 until her assassination in 1984. Her son, Rajiv Ghandi, was also a prime minister and he served following his mother's death until 1989. (He was later assassinated in 1991.) Since the Indira Awas Yojana was launched under Rajiv's administration and was named for his mother, this program was most likely named after him.
In 2009, as part of an effort to promote “slum-free cities,” the Indian Ministry of Housing and Urban Poverty Alleviation (HUPA) announced a new initiative, the Rajiv Awas Yojana (RAY).[1] The RAY aims to upgrade and bring all slums within the formal housing system. It also intends to prevent the development of new slums by providing more affordable housing options and developing more land to meet growing needs. As a prerequisite to RAY funding, states must first assign legal title to slum-dwellers over their living space.
In theory and on paper, the RAY includes a number of positive measures for slum-dwellers, especially the emphasis on property rights and the lack of distinction between government-recognized and unrecognized slums. If implemented properly, it could result in improved housing and access to services for many more of the urban poor.
However, in India there is a large disconnect between policy and implementation. The RAY Guidelines for Slum-free City Planning lack appropriate measures to ensure that slum-dwellers’ interests will be protected against arbitrary eviction and relocation, especially when private builders are involved in slum re-development. Without strict provisions for community participation, transparency, monitoring, and accountability, there is a danger that the RAY could be used to grab valuable slum lands in the center of the city while relocating poorer residents to faraway locations.
Potential problems with the RAY can be broken down into several categories:
1) Lack of Monitoring Mechanisms and Transparency
- States must develop a Plan of Action (POA) that follows central government RAY guidelines in order to receive funding. The reality? Inadequate plans are approved and since the POAs are not statutory documents, they are legally non-binding on states.
- Transparency is not emphasized in the functioning of the RAY. For example, maps and survey information about slums are not mandated to be publicly available, nor is information such as detailed project reports, lists of beneficiaries, consultants involved, and names of members of RAY governing and planning committees. A similar lack of transparency has led to corruption in the Slum Rehabilitation Authority in Mumbai.
- Frequency: the RAY does not stipulate with what frequency data should be updated, and in what timeframe measurable slum improvements should be made. Because slums are dynamic, living communities, regular surveys are required to maintain data that is useful for policy and planning.
- Limits of Geographic Information Systems (GIS): Using satellite imagery alone to identify slums may result in smaller slums going unnoticed, particularly in cities like Chennai, which have a large number of smaller slum clusters rather than a few large slum clusters.
- Lack of participation in data collection: Although community participation is recommended for data collection in the POA guidelines, it is not mandatory. Already, in Chennai, surveys have begun, but it does not seem that non-governmental organizations /community-based organizations or individual slum-dwellers have been involved in collecting survey information.
- The RAY guidelines stipulate that the POA should emphasize Public-Private-Partnerships (PPP) in slum redevelopment. Private players implementing RAY will be allowed to make commercial use of some areas or sell a few flats at market rates. Learning from the case of Mumbai’s Slum Rehabilitation Authority, there is a particular need for institutional safeguards to prevent the mismanagement of land use and to protect the rights of vulnerable slum-dwellers.
- Thus far, it is unclear how the RAY will be unified. In the state of Tamil Nadu, the state Slum Clearance Board (TNSCB) has taken the planning lead but how the TNSCB will work with other agencies with the Chennai Metropolitan Development Authority, the Town and Country Planning Department, the Tamil Nadu Housing Board, community groups, and the new “slum-free city technical cells” is unclear. Furthermore, once a Property Rights to Slum-Dwellers Act is passed, there will be additional agencies to coordinate around the question of land tenure.
- GIS-based slum planning, planning to prevent slums, and providing security of tenure to the poor are all complex tasks for which cities and housing agencies have little experience. Even with some money for technical consultants, it is unclear whether municipalities will have enough technical capacity to implement the RAY.
---------------
Notes
[1] While there is no exact translation of "Rajiv Awas Yojana," in Hindi "yojana" means "scheme," which in India is the word used for government programs or policies. "Awas" is commonly used for yojanas dealing with housing related issues. While Rajiv literally translates as "lotus flower," we think the term refers to something else. In India there is also a program called Indira Awas Yojana, started in 1985, to provide housing for the rural poor. Indira was the name of a female prime minister of India named Indira Gandhi, who served from 1966-1977 and from 1980 until her assassination in 1984. Her son, Rajiv Ghandi, was also a prime minister and he served following his mother's death until 1989. (He was later assassinated in 1991.) Since the Indira Awas Yojana was launched under Rajiv's administration and was named for his mother, this program was most likely named after him.
Friday, September 23, 2011
"Randomized control trials" on trial: Evaluating the efficacy of RCTs
Jake Velker, MPA
Are randomized trials the way to finally start making a dent in reducing poverty, after years of hopeful thinking and disappointing results? Does this tool for evidence-based policymaking hold the key for practitioners to determine which poverty reduction programs work and which don’t? These questions motivate two recent books published by researchers who are at the vanguard of the randomized control trials (RCT) movement: More Than Good Intentions by Dean Karlan and Jacob Appel of Innovations for Poverty Action (IPA) and Poor Economics by Abhijit Banerjee and Esther Duflo of the Abdul Latif Jameel Poverty Action Lab (J-PAL). They deliberately introduce randomization in the implementation of anti-poverty measures to provide confidence in the programs’ efficacy (or lack thereof).
The results thus far have been dramatic and not always intuitive: microfinance is less effective than we hoped [1]; free bed nets are used more often (and prevent more malaria) than those that cost money.[2] IPA and J-PAL are currently involved in dozens of trials and are generally credited with bringing an unprecedented level of rigor to the evaluation of development—a field normally dominated by grand theories and polemics.
Enough praise has been heaped on the “randomistas” that I feel confident I can focus on the criticisms of their methodology without sounding uncharitable.[3]
The first criticism—leveled forcefully by Princeton’s Angus Deaton—is that randomized trials do not help us in any systematic way to gain an understanding of why interventions work.[4] In this sense, IPA and J-PAL are part of a broader trend in economic research that eschews theory in favor of real-world applications and problem-solving. This is irksome for many economists, particularly those who believe that poverty cannot be solved without a broader accounting of the mechanisms that keep people trapped in poverty. Randomized trials are beginning to test theoretical frameworks more directly, but there is still much progress to be made.
The most relevant criticism, however, is political. IPA and J-PAL economists have been accused of ignoring the institutional constraints against which their interventions would inevitably contend if scaled up. Often, research from randomized trials offers a conclusion like “we find that intervention X lowers Y disease transmission by Z percent.” While it is extremely helpful to have confidence in the efficacy of a treatment, glaring questions remain. Does the program work when it is implemented by a weak bureaucracy, rather than university-trained researchers? At scale, who will be responsible for administering the recommended program? What are their incentives to perform? Where will the money come from? If the intervention is so successful, were there good reasons it wasn’t tried before?
A troubling case in point involves one of the most heralded studies from the RCT movement to date. Working in Kenya, economists Michael Kremer and Ted Miguel found that providing de-worming medicine to students boosted school attendance cost-effectively.[5] Spurred by their research, the Kenyan government committed to making de-worming medicine available to more than 3,000,000 of its primary school children in 2009. But the policy was recently discontinued due to a dispute between the Kenyan government and international donors over corruption and the administration of education funding.[6] It should go without saying that for the ultra-poor, these sorts of bureaucratic obstacles are the norm, rather than the exception.
If economics is just supply and demand, the work of IPA and J-PAL has focused thus far mostly on demand. It is difficult to quantitatively study governance—imagine what an RCT studying a poor country’s provincial governance, for example, might look like—and even harder to actually improve the quality of basic services in developing countries. So many of the interventions RCTs have found to be effective involve classic public goods, which by definition remain under-provisioned by private markets. But the bureaucracies of developing countries are generally ineffective, if not downright corrupt. This is where economics loses its relevance and institutions and leadership rear their ugly heads.
These problems have not been amenable to ever-more creative randomized trials. In fact, many of the most celebrated finds of the RCT movement are relative “no-brainers.” Who, after all, would argue against treating poor school children for intestinal worms? Esther Duflo and her colleagues have said that we do not know what works. Many would respond that we know perfectly well what works; but do not know how to do it. Perhaps the real questions start once an intervention has been proven to work.
Randomistas respond to this critique as follows. First, it was never their ambition to overhaul the political economy of the developing world. The fact that they have found real evidence of effective interventions is in itself a major accomplishment. They believe that their approach can improve lives even in discouraging political settings. They are not promising a sweeping social revolution, but rather a “quiet revolution” of incremental gains. And even critics will concede that though the modesty of this approach may be unsatisfying, it is nonetheless an improvement on the empty promises all too frequent in the development world.
------------------------
References
[1] Abhijit Banerjeey, Esther Duflo, Rachel Glennerster, and Cynthia Kinnan, “The miracle of microfinance? Evidence from a randomized evaluation,” Working Paper (unpublished), May 2009.
[2] Jessica Cohen and Pascaline Dupas, “Free Distribution or Cost-Sharing? Evidence from a Randomized Malaria Prevention Experiment,” Quarterly Journal of Economics, Vol. 125:1, 2010.
[3] For examples of such praise, see: Ian Parker, “The Poverty Lab: Transforming development economics, one experiment at a time,” New Yorker, May 2010; James Crabtree, “Attested Development,” Financial Times, April 2011; William Easterly, “Measuring How and Why Aid Works – or Doesn’t,” Wall Street Journal, April 2011; Ben Goldacre, “How can you tell if a policy is working? Run a trial,” The Guardian, May 2011; and Nicholas Kristof, “Getting Smart on Aid,” New York Times, May 2011.
[4] Angus Deaton, “Instruments, Randomization, and Learning about Development,” Journal of Economic Literature, Vol. 48:2, June 2010.
[5] Edward Miguel and Michael Kremer, “Worms: Identifying Impacts on Education and Health in the Presence of Treatment Externalities,” Econometrica, Vol. 72: 1, January 2004.
[6] Justin Sandefur, “Held Hostage: Funding for a Proven Success in Global Development on Hold in Kenya,” Global Development: Views from the Center blog, Center for Global Development, April 2011.
Are randomized trials the way to finally start making a dent in reducing poverty, after years of hopeful thinking and disappointing results? Does this tool for evidence-based policymaking hold the key for practitioners to determine which poverty reduction programs work and which don’t? These questions motivate two recent books published by researchers who are at the vanguard of the randomized control trials (RCT) movement: More Than Good Intentions by Dean Karlan and Jacob Appel of Innovations for Poverty Action (IPA) and Poor Economics by Abhijit Banerjee and Esther Duflo of the Abdul Latif Jameel Poverty Action Lab (J-PAL). They deliberately introduce randomization in the implementation of anti-poverty measures to provide confidence in the programs’ efficacy (or lack thereof).
The results thus far have been dramatic and not always intuitive: microfinance is less effective than we hoped [1]; free bed nets are used more often (and prevent more malaria) than those that cost money.[2] IPA and J-PAL are currently involved in dozens of trials and are generally credited with bringing an unprecedented level of rigor to the evaluation of development—a field normally dominated by grand theories and polemics.
Enough praise has been heaped on the “randomistas” that I feel confident I can focus on the criticisms of their methodology without sounding uncharitable.[3]
The first criticism—leveled forcefully by Princeton’s Angus Deaton—is that randomized trials do not help us in any systematic way to gain an understanding of why interventions work.[4] In this sense, IPA and J-PAL are part of a broader trend in economic research that eschews theory in favor of real-world applications and problem-solving. This is irksome for many economists, particularly those who believe that poverty cannot be solved without a broader accounting of the mechanisms that keep people trapped in poverty. Randomized trials are beginning to test theoretical frameworks more directly, but there is still much progress to be made.
The most relevant criticism, however, is political. IPA and J-PAL economists have been accused of ignoring the institutional constraints against which their interventions would inevitably contend if scaled up. Often, research from randomized trials offers a conclusion like “we find that intervention X lowers Y disease transmission by Z percent.” While it is extremely helpful to have confidence in the efficacy of a treatment, glaring questions remain. Does the program work when it is implemented by a weak bureaucracy, rather than university-trained researchers? At scale, who will be responsible for administering the recommended program? What are their incentives to perform? Where will the money come from? If the intervention is so successful, were there good reasons it wasn’t tried before?
A troubling case in point involves one of the most heralded studies from the RCT movement to date. Working in Kenya, economists Michael Kremer and Ted Miguel found that providing de-worming medicine to students boosted school attendance cost-effectively.[5] Spurred by their research, the Kenyan government committed to making de-worming medicine available to more than 3,000,000 of its primary school children in 2009. But the policy was recently discontinued due to a dispute between the Kenyan government and international donors over corruption and the administration of education funding.[6] It should go without saying that for the ultra-poor, these sorts of bureaucratic obstacles are the norm, rather than the exception.
If economics is just supply and demand, the work of IPA and J-PAL has focused thus far mostly on demand. It is difficult to quantitatively study governance—imagine what an RCT studying a poor country’s provincial governance, for example, might look like—and even harder to actually improve the quality of basic services in developing countries. So many of the interventions RCTs have found to be effective involve classic public goods, which by definition remain under-provisioned by private markets. But the bureaucracies of developing countries are generally ineffective, if not downright corrupt. This is where economics loses its relevance and institutions and leadership rear their ugly heads.
These problems have not been amenable to ever-more creative randomized trials. In fact, many of the most celebrated finds of the RCT movement are relative “no-brainers.” Who, after all, would argue against treating poor school children for intestinal worms? Esther Duflo and her colleagues have said that we do not know what works. Many would respond that we know perfectly well what works; but do not know how to do it. Perhaps the real questions start once an intervention has been proven to work.
Randomistas respond to this critique as follows. First, it was never their ambition to overhaul the political economy of the developing world. The fact that they have found real evidence of effective interventions is in itself a major accomplishment. They believe that their approach can improve lives even in discouraging political settings. They are not promising a sweeping social revolution, but rather a “quiet revolution” of incremental gains. And even critics will concede that though the modesty of this approach may be unsatisfying, it is nonetheless an improvement on the empty promises all too frequent in the development world.
------------------------
References
[1] Abhijit Banerjeey, Esther Duflo, Rachel Glennerster, and Cynthia Kinnan, “The miracle of microfinance? Evidence from a randomized evaluation,” Working Paper (unpublished), May 2009.
[2] Jessica Cohen and Pascaline Dupas, “Free Distribution or Cost-Sharing? Evidence from a Randomized Malaria Prevention Experiment,” Quarterly Journal of Economics, Vol. 125:1, 2010.
[3] For examples of such praise, see: Ian Parker, “The Poverty Lab: Transforming development economics, one experiment at a time,” New Yorker, May 2010; James Crabtree, “Attested Development,” Financial Times, April 2011; William Easterly, “Measuring How and Why Aid Works – or Doesn’t,” Wall Street Journal, April 2011; Ben Goldacre, “How can you tell if a policy is working? Run a trial,” The Guardian, May 2011; and Nicholas Kristof, “Getting Smart on Aid,” New York Times, May 2011.
[4] Angus Deaton, “Instruments, Randomization, and Learning about Development,” Journal of Economic Literature, Vol. 48:2, June 2010.
[5] Edward Miguel and Michael Kremer, “Worms: Identifying Impacts on Education and Health in the Presence of Treatment Externalities,” Econometrica, Vol. 72: 1, January 2004.
[6] Justin Sandefur, “Held Hostage: Funding for a Proven Success in Global Development on Hold in Kenya,” Global Development: Views from the Center blog, Center for Global Development, April 2011.
Friday, May 20, 2011
Farewell to Robertson: Thinking beyond metrics
Payal Hathi, MPA
I met with a woman this week that came into my office with huge bruises on her arms and face. She has been married to the man who gave her those bruises for 16 years, and they have two children together. She believes that she deserved what happened, that although she called the police for help in making the violence stop, she couldn’t actually tell them the truth about what was happening in her home; that if she just stays quiet and stays in the basement, her husband won’t get angry; and that all she needs to do is wait for 5 more years until her daughter is old enough to go to college. Then it will all be over.
Violence against women is not something that comes up often in policy circles – domestic violence in particular is often considered a family issue, or one that is generally addressed by activists. But there is certainly a place for policy to play a role. Societal portrayals and beliefs of women – from the media to childhood socialization to educational and job opportunities – all play a role in allowing the feminine to be seen as the inferior. In addition, the general attitude of blaming the victim and the cultural stigma often attached to women who “can’t keep their families together” perpetuates not only intergenerational cycles of violence, but also silence and shame around the issue.
A major part of the problem is that in the world of metrics and quantitative measures of success that we live in today, it often seems that there is little space or value given to work that cannot be counted using numbers. When I am asked how many women with whom I work have finally left their partners, found a job, or have secured housing for themselves, it seems much of the rest of the work that I do is invalidated because I cannot quantify what it means to get a woman to believe that she has the right to live free of fear or that she deserves to earn a living wage for her work. Of course the importance and power of statistics and clear, concise, and concrete information is critical in making policy issues relevant to people. But having worked in the violence against women movement for over a year now, it is clear that an additional system of evaluation is necessary if we as policymakers are to make the world a safer and more just place.
I came into graduate school with a background in quantitative evaluation work, hoping to gain the skills to enable small non-profits to incorporate evaluation as a routine practice to improve their own work. Through my work this past year with women facing human rights and economic justice violations in almost every aspect of their lives, I have become acutely aware that as policymakers, it is not enough to think about the efficiency with which different ideas can be put into practice, or the cost-effectiveness of one approach versus another. Rights, and the way in which they are put into practice in our families and daily interactions—while not easy to measure—must form the foundations of our policy frameworks in order to make meaningful impacts in the communities we serve. As I graduate, I recognize that as we look towards more sophisticated ways to evaluate the work of service providers, we cannot diminish the work that instills in people an awareness of their rights and gives them the tools to exercise them.
Payal Hathi worked at Sakhi for South Asian Women from 2010 to 2011. If you or someone you know has an issue with domestic violence, please go to Sakhi’s website http://sakhi.org or call their helpline at (21) 868-6741.
I met with a woman this week that came into my office with huge bruises on her arms and face. She has been married to the man who gave her those bruises for 16 years, and they have two children together. She believes that she deserved what happened, that although she called the police for help in making the violence stop, she couldn’t actually tell them the truth about what was happening in her home; that if she just stays quiet and stays in the basement, her husband won’t get angry; and that all she needs to do is wait for 5 more years until her daughter is old enough to go to college. Then it will all be over.
Violence against women is not something that comes up often in policy circles – domestic violence in particular is often considered a family issue, or one that is generally addressed by activists. But there is certainly a place for policy to play a role. Societal portrayals and beliefs of women – from the media to childhood socialization to educational and job opportunities – all play a role in allowing the feminine to be seen as the inferior. In addition, the general attitude of blaming the victim and the cultural stigma often attached to women who “can’t keep their families together” perpetuates not only intergenerational cycles of violence, but also silence and shame around the issue.
A major part of the problem is that in the world of metrics and quantitative measures of success that we live in today, it often seems that there is little space or value given to work that cannot be counted using numbers. When I am asked how many women with whom I work have finally left their partners, found a job, or have secured housing for themselves, it seems much of the rest of the work that I do is invalidated because I cannot quantify what it means to get a woman to believe that she has the right to live free of fear or that she deserves to earn a living wage for her work. Of course the importance and power of statistics and clear, concise, and concrete information is critical in making policy issues relevant to people. But having worked in the violence against women movement for over a year now, it is clear that an additional system of evaluation is necessary if we as policymakers are to make the world a safer and more just place.
I came into graduate school with a background in quantitative evaluation work, hoping to gain the skills to enable small non-profits to incorporate evaluation as a routine practice to improve their own work. Through my work this past year with women facing human rights and economic justice violations in almost every aspect of their lives, I have become acutely aware that as policymakers, it is not enough to think about the efficiency with which different ideas can be put into practice, or the cost-effectiveness of one approach versus another. Rights, and the way in which they are put into practice in our families and daily interactions—while not easy to measure—must form the foundations of our policy frameworks in order to make meaningful impacts in the communities we serve. As I graduate, I recognize that as we look towards more sophisticated ways to evaluate the work of service providers, we cannot diminish the work that instills in people an awareness of their rights and gives them the tools to exercise them.
Payal Hathi worked at Sakhi for South Asian Women from 2010 to 2011. If you or someone you know has an issue with domestic violence, please go to Sakhi’s website http://sakhi.org or call their helpline at (21) 868-6741.
Subscribe to:
Posts (Atom)