Showing posts with label Improving Psychology. Show all posts
Showing posts with label Improving Psychology. Show all posts

Saturday, May 31, 2025

Question: Can we rid science of retracted articles?

I saw this opinion article on the problem of zombie papers and some possible solutions (i.e., retracted papers that continue to get cited and have influence) thanks to Retraction Watch. You can read the original in French here. Marc Joets makes some useful points here. Yes, we definitely have a problem. Part of the problem is that it takes a long time on average to get a problematic article (i.e., one that has error-ridden data reporting, fraudulent data reporting, or plagiarism) retracted. Apparently from date of publication to date of retraction, you're probably expect three years before a problematic paper is retracted, give or take. Apparently there are regional differences in time taken to retract an article: American and western European based journals do so more rapidly than elsewhere. Subscription based journals tend to retract articles more rapidly once a problem is identified than open source journals. In other words we need to keep in mind that there may be variations in culture and editorial practices at play. 

I've mentioned retractions before, and zombie articles before. In our various scientific fields, zombies are a legitimately concerning problem. As long as they are cited, they risk infecting not only the specific scientific discipline in question but also public discourse and policy. If you are living in the US right now, you probably know that a fraudulent and retracted article that spread some outright lies about the safety of childhood vaccinations has led in a matter of a couple decades to mainstream an anti-vax movement that is now in control of our own federal public health agencies. In this case, the consequences are life and death as the government is no longer as interested in containing a deadly measles outbreak. In my corner of the scientific community, the stakes may be considerably lower, but zombie articles can still infect public discourse and policy in ways that are not in the public interest. 

So, what to do? The answers in this editorial are ones that strike me as common sense at this point. Making data and research protocols publicly available can help to catch mistakes and fraud early enough to nip the problem in the bud. Better plagiarism detection tools are mentioned as well. Ultimately the author notes that there is no one-size-fits-all solution. But in broad brushstrokes we can expect that efforts to beef up transparency help. Efforts to improve reproducibility - requiring pre-registration of research protocols and offering evidence of replicability -are also necessary. Any of these practices can help detect errors or problems in a more timely manner. In addition to better plagiarism detection tools, the author suggests that each journal have its own panel that can objectively handle instances of fraud or serious errors as they occur. Finally, the author argues for making the fact that these articles have been retracted more visible in order to minimize the impact of retracted work. That strikes me as a solid idea. I still get a sense that retractions are not nearly as visible as they could be. Those with the PubPeer browser extension might be a bit more wise to retractions, as are those who use the Retraction Watch site's own retraction database. But how many of us are actually using those resources currently and consistently? I wonder. 

I wish that instructions for eliminating zombies from our sciences were as simple as "destroy the brain or remove the head"* but alas they are not. I remain a cautious optimist however.

*The reference in that quote was specifically to "Shaun of the Dead" which is a personal favorite of mine, but probably would refer to most zombie films and series I've seen over the years. 

 

Monday, May 5, 2025

Replications and Reversals

A few years ago, I wrote about a blog post documenting reversals in psychology. In the intervening years, it has turned into a full-fledged site hosted by FORRT (Replications and Reversals), and is quite comprehensive. If you are interested in what classic research has held up and what has at best mixed evidence or has been debunked, this is a valuable resource. Every time I redo my social psychology course, this is a site that will give me more ammo to make sure that my students have the most accurate information available.

Tuesday, March 25, 2025

Back to Zimbardo's Stanford Prison Experiment

A few months ago, I saw an article on the Stanford Prison Experiment, discussed it on this blog in a brief post in which I said I wanted to circle back to this. Thankfully Retraction Watch gave me an excuse, and some down time at a conference is allowing me time to say a few more words. Zimbardo's Stanford Prison Experiment (SPD going forward) was controversial from the start, as any of us with even a cursory awareness of social psychological research would know. Regrettably, it has been taken as the gospel in many textbooks (Introductory Psych textbooks as well as Social Psych textbooks) leading to the impression that the SPD was instrumental in establishing the importance of the power of the situation. Zimbardo himself certainly used the notoriety of the SPD for self-promotion throughout the remainder of his career and life. 

The Retraction Watch article I mentioned above addresses a question that seems reasonable to ask: should Zimbardo's papers on SPD be retracted? Certainly, as the guest author notes, there are good reasons to do so. First up the author considers the scientific credentials of the study. There is no well-defined a-priori hypothesis which itself is a big no-no. Where was the control group? There was none. Any tests of significance? Who needs those, right? Then there is the little matter of Zimbardo's expertise. He had no known background in criminology as I understand it. One major ethical principle is competence. Does the PI have the background to adequately design and execute the study? Can the PI adequately train their research assistants to carry out their duties in the lab? In this case, Zimbardo fails the competency test.

There is also the question of the originality of the study's design. Granted, in the sciences we build on the work of others. So as far as novelty goes, I am not that much of a stickler. But it is important to credit the sources that inspired one's work. In the case of the SPD, apparently Zimbardo got the idea from a student term paper that remained uncredited.

The ethical treatment of participants is certainly worth weighing as well. The conditions the "prisoners" in this simulation experienced were inhumane to put it bluntly. Humiliation, unsanitary conditions, and sleep deprivation were all part of the experience. The SPD was a lawsuit waiting to happen. 

I have covered the credibility of the findings elsewhere on this blog. There was certainly evidence of fraud in the reporting. Important details were left out intentionally. The seemingly "spontaneous" scene during the six days the simulation ran were effectively stage managed by Zimbardo. Maybe I am being charitable calling this the psychology of community theatre, but it captures the essence.

So all of the above would make any paper published ripe for retraction. But there is a problem. This research is over 50 years old. One article that was published in the early 1970s appeared in a journal that technically no longer exists. Retraction would be impossible in that case. 

If retraction is impossible, what is the remedy? I think I am largely in agreement with the author that the SPD is best covered in textbooks, seminars, and public media coverage as a cautionary tale of how not to conduct research or how to report findings. The study had so many ethical red flags after all. And really this is my clarion call to textbook authors: if you are working on a new book for students or updating your own, and you cover Zimbardo's SPD, make sure to treat it not as some amazing revelation about human behavior but rather as a flawed and fraudulent study that violated fundamental ethical principles. Don't leave those of us who teach Introductory Psych courses or Social Psych courses do damage control individually. 

Monday, May 20, 2024

Interesting podcast on Open Science and its enemies

Now that I have a couple of moments to breathe, I've been able to spend a bit of time on Bluesky (which is where a lot of academic Twitter landed after Elon Musk took over the platform and made it far worse), and reconnect with some folks whose work I respect. I have been gathering that there might be some trouble in paradise among the community advocating for Open Science. I've seen some chatter about a preprint that offers a very broad definition of what may be considered questionable research practices that have at least some in the community suggesting that the definition really is too broad. I may come back around to that, but if and when I have the time to do more than give the preprint an initial reading. Instead, I'll focus on a podcast that I had never heard of before entitled The Error Bar. It's a clever title, and the host, Nick Holmes, certainly strikes me as witty and open-minded. He is planning a three-part series on what he refers to as Open Science and its Enemies. This week's podcast is called the p-circlers. We can think of p-circling as reverse p-hacking according to Dr. Holmes. The upshot is that p-circlers hone in on a finding they do not like and then look for ways to make the finding seem suspect or so trivial as to not even be worth examining in the first place. Holmes focuses on one finding that seemed to create a stir, and a preprint that ends up resorting to p-circling behavior in order to explain away the findings as much ado about nothing. If you have 26 minutes or so to spare, it is a worthwhile podcast episode, and hopefully will provoke some thought. Since I have some proverbial skin in the game when it comes to presenting open science practices in my undergraduate methods courses, I want to make sure that our actions really do move our respective disciplines and sub-disciplines forward, rather than simply weaponize a series of recently developed tools for post-peer-review and in turn lead us to making the same mistakes as our predecessors. Anyway, give this episode a listen. I don't think you'll be disappointed.

P.S.: If you do not like listening to podcasts, Nick Holmes also blogs his podcasts. Here is the blog post for the episode on p-circlers.

Saturday, July 10, 2021

I published a meta-analysis, and now there is a retracted study in my database. Should I worry?

Since the title addresses a very real question for me, it's worth asking, as a soon-to-be-retracted article was included in our weapons effect database. Fortunately, Fanelli, Wong, and Moher (2021) address this very question: what impact do retracted studies have on the conclusions we can draw from our meta-analytic findings? In other words, what are the epistemic costs? The good news is not much. The authors took a sample of 50 or so meta-analyses that had included at least one retracted study. The positives are really positive - findings tend to remain robust even after a retracted study or studies are removed. This is especially important to the extent that this finding holds when the retraction was due to something suspect in the methodology or in the data analyses, and not some other issue such as plagiarism. One thing that the authors do note as that much of the problem of retracted studies appearing in meta-analyses is preventable. Many of of the meta-analyses Fanelli et al. (2021) included in their meta-meta-analysis had included studies that had been retracted well before the meta-analyses in question were published. They have their ideas of some systemic corrections that would help. I would recommend including the PubPeer web browser plug-in as one means of screening for potential retractions early on in the meta-analytic database search process. It won't catch everything, especially to the extent that it is underutilized by psychologists, but it could help a bit. I would also recommend searching through the Retraction Watch database. Those are individual actions we can take, and take now. 

Hat tip to Retraction Watch.

Tuesday, June 29, 2021

Reversals in Psychology

This blog post captures the essence of something I've been increasingly imparting to my students, especially over the last three or four years. Like any other scientific field, we're going to have reversals. Something we thought might be true turns out to be not only falsifiable, but just plain false. The author is fairly optimistic about Psychology. We're a bit more transparent, if less than ideal, than other fields, and that has enabled us to call ourselves out on our own bovine fecal matter. If nothing else, it's a reminder that many of our classics, which are still taught in textbooks and portrayed in pop culture as true, are not necessarily what they appear to be on the surface.

Sunday, June 6, 2021

Is there a brain drain in the science reform movement?

The answer, according to Alexander Danvers, appears to be yes, there is indeed a brain drain. There are plenty of reasons why we appear to be losing our best and brightest, at a time when we arguably need them the most. There doesn't appear to be much of an incentive for reformers to do their work, including post-peer review necessary to weed out grossly incompetent and fraudulent research. Nor is there much of an incentive to develop or engage in the sort of necessary work of conducting replication research, developing and validating our measures in a way that would inspire confidence, etc. Certainly the grant money isn't there for such work. And gaining a reputation for engaging in reform-minded research activities is a terrible way to get promoted, given the way the power structure in the academic world currently works. There simply are not enough mid and late career scholars willing to defend this necessary work, and those who carry out that work. There's also the question of whether what we do in my field has much meaning. That's certainly a question that haunted Joe Hilgard as he contemplated his eventual exit from academic life. Indeed, one might make more of a difference as a data scientist in any of a number of industries. And although I am quite happy for my peers who have found more lucrative and rewarding careers outside of the academic world, I can't help but wonder how that bodes for the future of reform. How much of what some very driven and competent reformers within psychology will become normative? How much will get set aside as the publish-or-perish model of scholarly life continues to dominate, and those who have profited from the old status quo continue to call the proverbial shots? Could independent research centers like IGDORE Institute be a way of sidestepping at least some of the current power structure? In the meantime, on a more personal note, I am having to accept that at least some subset of the people I met at SIPS in 2019 and again virtually in 2020 are ones who will not be around once I can finance another international conference trip in a couple years. I'll miss them. Hopefully some newer members at SIPS will be ready to carry the torch further. We shall see. Whatever happens, we need to make sure that there are incentives in place to keep our best minds with us. Otherwise, my field is one that will deserve to slip into irrelevance.

Wednesday, April 7, 2021

Recent Vox article on the Psychological Science Accelerator

 Vox has been covering the replication crisis for some time. This time, Vox has an article on the Psychological Science Accelerator, its foundation, successes (so far), and some tangible challenges (including funding). This is worth a read, if for nothing else one can get an idea of what an approach to open science looks like in practice.

Friday, August 14, 2020

Uli Schimmack on a decade of replication failures

I am more familiar with Uli Schimmack from the Facebook group he runs (Psychological Methods Discussion Group) and his R-Index handle on Twitter. He also regularly blogs, and I have used some of his posts as supplemental materials in my undergraduate Social Psychology course. Earlier this year, he posted a reflection on the replication crisis that hit Social Psychology especially hard initially (I often say we were ground zero). I recommend the post, which has been updated to reflect the content of an article he published recently. Well worth your time.

Thursday, January 30, 2020

Another Stapel situation?

It's hard to say right now. What is clear is that another PI has seen a couple of his papers retracted due to data irregularities. His former students, post-docs, and coauthors are doing the right thing. At the end of the day, that's what matters. One thing to keep in mind: when we start looking at cases of potential fraud in scientific research, the people most affected are early career researchers (ECRs): grad students and post-docs in particular. With fewer lines on a CV, any retraction will have disproportionate repercussions any time they are on the job market, applying for tenure, etc. All of us, though are negatively affected, to the extent that any article based upon fraudulent (or fabricated) data, to the extent that policy,health, and other personal or professional decisions are based upon that research. The old Russian saying "trust but verify" is seeming more apt all the time. Maybe I'd leave the "trust" part out and just say "verify."

Friday, December 27, 2019

Data sleuthing made easy

You all know that I have done just a bit of data sleuthing here or there. I do so with no real fancy background in statistics. I have sufficient course work to teach stats courses at the undergraduate level, but I am no quantitative psychologist. So, I appreciate articles like How to Be a Statistical Detective. The author lays out some common problems and how any of us can use our already existing skills to detect those problems. I use some of these resources already, and am reasonably adept at using a calculator. I will likely add more links to these resources to this blog.

This article is behind a paywall, but I suspect my more enterprising readers already know how to obtain a copy. This article is fundamental reading.

Monday, November 4, 2019

Another resource for sleuths

This tweet by Elizabeth Bik is very useful:

The site she used to detect a publication that was self-plagiarized not only in terms of data and analyses but also in terms of text can be found here: Similarity Texter. I will be adding that site to this blog's links. I think as a peer reviewer it will help in detecting potential problem documents. Obviously I see the utility for post-peer review. Finally, any of us as authors who publish multiple articles and chapters on the same topic would do well to run our manuscripts through this particular website prior to submission to any publishing portal. Let's be real and accept that the major publishing houses are very lax when it comes to screening for potential duplicate publication, in spite of the enormous profits that they make from taxpayers across the planet. We should also be real about the quality of peer review. As someone who has been horrified to receive feedback on manuscript from a supposedly reputable journal in less than 48 hours, I think a good case can be made as an author for taking things into your own hands as much as possible. That along with statcheck can save some embarrassment as well as ensure that we as researchers and authors do due diligence to serve the public good.

Sunday, August 25, 2019

Will this time be different?

I had the pleasure of seeing and hearing Sanjay speak these words live in early July at the closing of the SIPS conference in Rotterdam. I really hope those words are heeded. Simply tightening up some methods without addressing the social inequality that afflicts our science (as is the case with so many sciences) is insufficient. If the only people who benefit are those who just happen to keep paying membership dues, we've failed. Open science is intersectional and is a social movement. Anything short of that will be a failure. I hope that I do not find myself in a decade asking the same question that a good friend of mine once asked over three decades ago in his zine, Pressure: "So, where's the change?"

Thursday, August 1, 2019

Some initial impressions about SIPS 2019

I think perhaps the best way to start is with a Twitter thread I posted right as we were about to end:


The conference was different from any conference I have ever experienced. For those wanting to get a feel for what SIPS is about, a good place to begin might be to check out the page for this year's conference. Rotterdam was a good location in part because the city slogan is make it happen. SIPS is an organization devoted to actively changing the way the science of psychology is done, and is formatted in such a way that those participating become active. This is a conference for people who really want to roll up their sleeves and get involved.

My experience started with the preconference put together by the repliCATS project. Their travel grant to those willing to participate in the preconference is what made going to SIPS possible for me. During the 5th and 6th of July, I spent the entire work day at the conference site with a team of several other psychologists in various phases of their careers (most were postdocs and grad students). Each team was tasked with the responsibility of assessing the probability of replication for 25 claims. We had a certain amount of time to read each claim, look up the relevant article, look up any other supplementary materials relevant to the task, and then to make our predictions. We then discussed our initial assessments and recalibrated. I found the process engaging and enlightening. What was cool was how each of us brought some unique expertise to the table. Four of the claims we assessed were meta-analytic, and since I was the one person in the room who had conducted published meta-analyses, I became the de facto expert on that methodology. Believe me when I say that I still don't feel like an expert. But okay. So for those claims, my peers quizzed me a good deal and I got to share a few things that might be useful about raw effect sizes, assessment of publication bias, etc. We did the same with others. There were some claims that all of us found vexing. Goes with the territory. I think what I got from the experience was how much we in the psychological sciences really need each other if we are going to move the field forward. I also realized just how talented the current generation of early career researchers truly is. As deflating as some of the research we evaluated was, I could not help but feel a certain level of cautious optimism about the future of the psychological sciences.

The conference itself started mid-day on the 7th of July and lasted through the 9th of July. It was filled with all sorts of sessions - hackathons, workshops, and unconferences. I tended to gravitate to the latter two classes of sessions during the conference this time around. The conference format itself is rather loose. One could start a session, decide "no this is not for me" and leave without worrying about hurt feelings. One could walk in to a session in the middle and be welcomed. It was great to go to a professional meeting without seeing one person wearing business attire. At least that level of pretense was dispensed with, for which I am grateful. I gravitated toward sessions on the last day that had a specific focus on inclusiveness. That is a topic that has been on my mind going back to my student activist days in the 1980s. There is a legitimate concern that by the time all is said and done, we'll manage to fix some methodological problems that are genuinely troubling for the discipline without addressing the problem that there are a lot of talented people who could offer so much who are shut out due to their ethnicity, national origin (especially if from the Global South), sexual orientation, and gender identity. If all we get as the same power structure with somewhat better methods, we will have failed, in my personal and professional opinion. I think the people leading at least a couple of those sessions seemed to get that. I hope that those who are running SIPS get it too. Maybe an inclusiveness hackathon is in order? I guess I am volunteering myself to lead that one just by blogging about it!

This was also an interesting conference in that I am both mid-career and primarily an educator. So, I was definitely part of a small subset of attendees. Personally, I felt pretty engaged. I can see how one in similar circumstances might end up feeling legitimately left out. There was some effort to have sessions devoted to mid-career researchers. We may want that expanded to mid-career educators as well. We too may want to be active participants in creating a better science of psychology, but our primary means of doing so is going to be in the classroom and not via published reports. Those of us who may one day become part of administration (dept chairs, deans, etc.) or who already occupy those position should have some forum to discuss how we can better educate the next generation of Psychology majors at the undergraduate level so that they are both better consumers of research and better prepared for the changes occurring that will impact them as they enter graduate programs (for that subset of majors who will do so).

It was also great to finally meet up in person with a number of people with whom I have interacted via Twitter DM, email, and sometimes via phone. That experience was beautiful. There is something about actually getting to interact in person that is truly irreplaceable.

Given the attendance at the conference, I can see how much of a logistical challenge the organizers faced. There were moments where last minute room changes did not quite get communicated. The dinner was one where participants were mostly underfed (that is probably more on the restaurant than the organizers, and I can chalk that up to "stuff happens" and leave it at that). Maintaining that sense of intimacy with a much larger than anticipated group was a challenge. But I never felt isolated or alone. There was always a session of interest. There was always someone to talk to. The sense of organized chaos is one that should be maintained.

I did find time to wander around the city of Rotterdam, in spite of my relative lack of down time for this conference. The fact that it was barely past the Summer Solstice meant that I got some good daylight quality photos well into the evening. I got to know the "cool district" of Rotterdam quite well, and definitely went off the beaten path in the process. I'll share some of those observations at another time.

Overall, this was a great experience. I will likely participate in the future, as I am finding ways of using what I learned in the classroom. As long as participants can come away from the experience thinking and knowing that there were more good sessions than they could possibly attend, it will continue to succeed. If the organizers are serious about inclusiveness, they will have something truly revolutionary as part of their legacy. Overall, the sum total of this set of experiences left me cautiously optimistic in a way that I have not been in a very long time. There is hope yet for the psychological sciences, and I got to meet some of the people who are providing the reason for that hope. Perhaps I will meet others who did not attend this year at future conferences. I'll hold out hope for that as well.

Tuesday, May 28, 2019

The Clampdown: Or How Not to Handle a Crisis of Confidence

Interesting post by Gelman recently. I am on the mailing list from where the quoted email came. It was in reference to revelations about the Zimbardo prison experiment that cast further doubt on its legitimacy. As someone watching HBO's Chernobyl series, there is something almost Soviet in the mindset expressed in that email. The thing about clampdowns is that they tend to generate further cynicism that erodes the edifice upon which a particular discipline or sub-discipline is based. If I could, I'd tell these folks that they are just making the changes they are fighting more inevitable, even if for a brief spell the life of those labeled as dissidents is made a bit more inconvenient.

The title for this and Gelman's post is inspired by a song by The Clash:

Wednesday, May 1, 2019

Correcting the scientific record does take a toll on the people involved

May Day was once an international labor day. In the US we tend to avoid so much as a mere acknowledgement of this date and its meaning. It is one worth remembering.

In that spirit, I would like to offer a recent blog post by James Heathers. Part of the labor involved in the sciences includes cleaning up the messes that fellow scientists leave behind when they are grossly incompetent or commit outright fraud. That labor is rarely recognized, is often demonized, and is often done after hours in relative isolation. I don't see that as a sustainable way for individual scientists to live, nor is this a sustainable model for the sciences if we wish them to remain legitimate.

This is as good a time as any to question the incentive structure that allows poor quality work and fraudulent work to slip through the cracks. This is as good a time as any to challenge the publish or perish ethos that has come to define institutions that are not even traditionally research-oriented. The problems revealed by the replication crisis in my field are international in scope and require international solutions. We need to have funding in place for those who are doing the necessary cleanup work as well as structures in place that reward openness - especially when inevitable mistakes get made. We are so far away from where we need to be. We should not see people who are genuinely trying to do the right thing burn out. They need collective support, ASAP.

A wiser man than me once said something about "having nothing to lose but your chains." I think there may be something to that - especially for those of us who toil in relative obscurity and with relative job insecurity.

Thursday, April 4, 2019

What attracted me to those who are trying to reform the psychological sciences?

That is a question I ask myself quite a bit. Actually the answer is fairly mundane. As the saying goes" "I've seen stuff. I don't recommend it."

Part of the lived reality of working on a meta-analysis (or the three I have worked on) is that you end up going down many of your research area's dark alleys. You see things that frankly cannot be unseen. What is even more jarring is how recently some of those dark alleys were constructed. I've seen it all: overgeneralization based on small samples relying on mild operational definitions of the variables under consideration, poorly validated measures, the mere fact that many studies are inadequately powered, and so on.

If you ever wonder why phenomena do not replicate, just wander down a few of our own dark alleys and you will understand rather quickly. The meta-analysis on the weapons effect, for which I was the lead author, was a huge turning point for me. Between the allegiance effects, the underpowered research, and some questions about how any of the measures employed were actually validated, I ended up with questions that had no satisfactory answer. I've been able to show that the decline effect that others had found when examining Type A Behavior Pattern and health outcomes also applied to aggressive behavioral outcomes. I was not surprised - only disappointed in the quality of the research conducted. That much of the research was conducted at a point in time in which there were already serious questions about the validity of what we refer to as Type A personality is itself rather disappointing. And yet, that work persisted for a while, often with small samples. I have also documented in my last two meta-analyses duplicate publications. Yes, the same data sets manage to appear in at least two different journals. I have questions. Regrettably, those who could answer are long since retired, if not deceased. Conduct a meta-analysis, and expect to find ethical breaches, ranging from potential questionable research practices to outright fraud.

That's a long way of saying that I get the need for doing whatever can be done to make what we do in the psychological sciences better: validated instruments, registered protocols and analysis plans, proper power analyses, and so on. There are many who are counting on us getting it as close to right as is humanly possible. Those include not only students, but the citizens who fund our work. There is no point to "giving away the science of psychology in the public interest (as George Miller would have put it) if we are not doing due diligence at the planning phase of our work.

Asking early research professionals to shoulder the burden is unfair. Those who are in more privileged positions need to step up. We need to be willing to speak truth to the teeth of power, otherwise there is no point in us even continuing, as all we have is a pretense with little substance. I wish I could say doing so would make one more marketable and so on. The reality is far more stark. At minimum we need to go to work knowing we have a clean conscience. Doing so will maintain public trust in our work. Failure is not something I even want to contemplate.

So I am a reformer. However long I am around in an academic environment, that is my primary role. Wherever I can support those who do the heavy lifting, I must do so. I have undergraduate students and members of the public in my community counting on it. In reality, we all do.

Monday, April 1, 2019

When social narratives and historical/empirical facts clash

This is just intended to be a quick fly-by post, inspired by a talk an anthropologist friend of mine gave at a local bookstore over the weekend. In discussing the tourist industry centered on the frontier mythology that dominates my community, he noted that there were social facts that often hid historical facts. By the way, countering the prevailing narrative is not an easy task, and a great way to make a few enemies along the way.

Anyway, his presentation got me thinking a good deal about how we go about teaching the psychological sciences, and more specifically social psychology. Any of us who have ever taken an introductory psychology course will inevitably read and be lectured on the story of Kitty Genovese, who was murdered in Queens in the early 1960s. I won't repeat the story here, but I will note that what we often see portrayed in textbooks is not quite what actually happened. The coverage spawned research on the Bystander Effect, which may or may not be replicable. As a narrative, Genevese's murder has been used by social conservatives as an example of the breakdown of traditional moral values in modern society (part of the subtext is that Ms. Genovese was a lesbian), and by social psychologists to further their narrative of the potentially overwhelming power of the situation. Policymakers have used the early Bystander Effect research based upon the myth of Kitty Genovese to pass "Good Samaritan" laws. The Bystander Effect and the myth of Kitty Genovese that spawned that research has been monetized by ABC in the reality series, What Would You Do. There is power and profit to be had by maintaining the narrative while burying the historical facts. After several decades, the damage is done. The APA's coverage of the murder of Kitty Genovese effectively debunked the myth a decade ago. And yet it still persists. If more non-replications of the Bystander Effect across nations and cultures are reported, that is great for science - to the extent that truth is great for science. However, my guess is that the debunked classic work will remain part of the narrative, shared on social media, and in textbooks for the foreseeable future.

In my own little corner of research, there is a sort of social narrative that has taken hold regarding various stimuli that are supposed to influence aggressive and violent behavior. It is taken as a given that media violence is causally related to aggression and even violence, even though skeptics have successfully countered the narrative with ample data to the contrary. Even something as superficial as short-term exposure to a gun or a knife is supposed to lead to aggressive behavior. The classic Berkowitz and LePage (1967) experiment is portrayed in social psychology textbooks as a prime example of how guns can trigger aggressive behavioral responses. Now lab experiments like the one Berkowitz and LePage conducted are often very artificial, and hence hard to believe are real. But what if I were to tell you that some researchers went out into the field and found the same effect? You'd be dazzled, right? Turner and his colleagues (1975) ran a series of field experiments that involved drivers blocking other drivers at intersections. They measured whether or not a horn was honked as their measure of aggression. After all, a horn honk is loud and annoying, and often in urban environments is used as the next best thing to screaming at other drivers. At least that is the thinking. Sometimes the driver (a member of the research team) drove a vehicle with a rifle on a gun rack. Other times the driver did not. The story goes that Turner and colleagues found that when the gun was present, the blocked drivers honked their horns. Case closed. The problem was, Turner and colleagues never actually found what we present in textbooks. Except for a possible subsample - males driving late model vehicles - for most drivers the sight of a firearm actually suppressed horn honking! That actually makes sense. If you are behind some jerk at a green light who has a firearm visibly displayed, honking at them is a great way to become a winner of a Darwin Award! What was actually happening then, is that with the possible exception of privileged males, drivers tended to make the correct assessment that there was a potential threat in their vicinity and that they should act cautiously. For the record, as far as I am aware, after the Turner and colleagues report was published, no one has been able to find support that the mere presence of a gun elicits horn honking. And yet the false social narrative continues to perpetuate. Who benefits? I honestly don't know. What I do know that I see the narrative of the weapons effect (or weapons priming effect) used by those who want to advocate for stricter gun laws - a position I tend to agree with, although the weapons effect as a body of research is probably very ineffective as a means of building an argument. Those who benefit from censoring mass media may benefit again from the power they accumulate. Heck, enough lobbying got rid of gun emojis on iPhones a few years ago, even though there is scant evidence that such emojis have any real impact on real world aggression, let alone violence.

Finally, I am reminded of something from my youth. When I was a teen, the PBS series Cosmos was aired. I had read Sagan's book The Cosmic Connection just prior, and of course was dazzled by the series, and eventually by the book. I still think it is worth reading, though with a caveat. Sagan tells the story of Hypatia, a scientist that probably any contemporary girl would want to look up to, and her demise. As the story goes, Hypatia was the victim of a gruesome assassination incited by a Bishop in Alexandria (now in modern day Egypt), and that the religious extremists of the time subsequently burned down the Library of Alexandria. Eventually I would do some further reading and realize that Hypatia's apparent assassination was much more complex of a story than the one Sagan told, and that the demise of the Library of Alexandria (which was truly a state-of-the-art research center for its time) was one that occurred over the course of centuries. Sagan's tale was one of how mindless fanaticism destroyed knowledge. It's a narrative that I am quite sympathetic toward. And yet, the tale is probably not quite accurate. The details surrounding Hypatia's murder are still debated by historians. The Library's demise is one that can be attributed to multiple causes, including government neglect, as well as the ravages of several wars.

Popular social narratives may play on confirmation bias - a phenomenon any of us is prone to experiencing - but the historical or empirical record may tell another story altogether. If the lessons from a story seem too good to be true, they probably are. A healthy dose of skepticism is advised, even if not particularly popular. In the behavioral and social sciences, we are supposed to be working toward finding approximations of the truth. We are not myth makers and story tellers. To the extent that we accept and perpetuate myths, we are doing little more than science fiction. If that is all we have to offer, we do not deserve the public's trust. I think we can do better, and often really do better.

Monday, February 18, 2019

A note of caution in justifying the use of the Aggressive Word Completion Task

The following blog post was co-authored by James Benjamin at University of Arkansas-Fort Smith and Randy McCarthy at Northern Illinois University.

Following the larger trends in social psychology, the past few decades in aggression research have been heavily influenced by theories examining the cognitive structures and cognitive processes that contribute to aggressive behavior. Accordingly, researchers have developed several tasks to measure “aggressive cognitions.” Good measures of aggressive cognitions (i.e., those that are reliable and valid) allow researchers to test these social-cognitive theories; bad measures (i.e., unreliable or those lacking validity evidence) do not allow researchers to test these social-cognitive theories.

Recently, we have been looking at the published literature of one commonly-used task for measuring aggressive cognitions: The Aggressive Word Completion Task (AWCT).

The most commonly-used stimuli for this task are freely available here (and have been freely available much longer than OSF has existed), which we think is swell. The posted instructions for using these stimuli recommend that researchers cite Anderson et al. (2003), Anderson et al. (2004), and/or Carnagey and Anderson (2005). After closely examining these three articles, we recommend that researchers NOT cite these articles as evidence for the validity of the AWCT.

What is the AWCT?

When completing the AWCT, participants are presented with several word fragments (e.g., words with one or more missing letters) and are instructed to fill in the missing letters to create a word. Critically, some of the word fragments can be completed with a word that is either semantically associated with the concept of aggression or with a word that is not semantically associated with the concept of aggression. For example, the word fragment “KI_ _” can be completed with aggressive words, such as “KILL” or “KICK”, or it can be completed with non-aggressive words, such as “KITE” or “KISS.” Computationally, the more potentially aggressive word fragments that are completed with aggressive words is considered to indicate more aggressive cognitions.

Thus, the AWCT has several desirable features. First, the AWCT does not require expensive equipment or software. Second, the AWCT can be administered in pencil-paper or online formats. Third, the AWCT is typically described as a task that is widely applicable to a range of research questions. These advantages make the AWCT an easy and flexible task to use.

What is the validity evidence in Anderson et al. (2003), Anderson et al. (2004), and Carnagey and Anderson (2005)?

The earliest published study that used the AWCT is Anderson et al. (2003). Anderson et al. (2003) cites Anderson et al. (2002) as a justification for using the AWCT, which is a manuscript described as “submitted for publication.” Anderson et al. (2002) is eventually published and becomes Anderson et al. (2004). Anderson et al. (2004; which was previously cited as Anderson et al., 2002) then cites Anderson et al. (2003) as justification for using the AWCT.




Right off the bat you can see that Anderson et al. (2003) and Anderson et al. (2004) are essentially two studies using one another in a sort of Penrose staircase of justification for using the AWCT.




By 2005, Carnagey and Anderson (2005) describes the AWCT as a “valid measure of aggressive cognitions” and jointly cites Anderson et al. (2003) and Anderson et al. (2004) as justification.





To our eye, there is no independent validation work that is publicly available (which is not to say that any validation work was not done, just that it is not available). These early studies both start using the AWCT and point to themselves as justification for doing so.

The lack of independent validation work is not great news, but it also isn’t inherently damning. The implication is the data from these studies provide some evidence for the validity of the AWCT because these studies purportedly demonstrate the AWCT changes predictably to theoretically-relevant conditions. Indeed, in Anderson et al. (2003), Experiment 4, participants’ trait hostility was associated with their AWCT performance (F [1, 134] = 4.21, p = .042) and participants who were exposed to a song with violent lyrics had higher AWCT performance than two other comparison conditions (F [2, 134] = 3.26, p = .073; note that the authors reported this p-value as being less than .05, which means that either the reported F ratio is incorrect or the reported p-value is incorrect). In Anderson et al. (2003), Experiment 5, participants who were exposed to a song with violent lyrics had higher AWCT performance than those who were exposed to songs without violent lyrics (F [1, 141] = 6.16, p = .014). In Anderson et al. (2004), participants who played a violent video game used more aggressive word completions than participants who played a non-violent video game (F [1, 120] = 4.26, p = .041). And in Carnagey and Anderson (2005), there were main effects for type of video game on aggressive word completions (F [2, 59] = 5.33, p = .007) and baseline systolic blood pressure (F [1, 59] = 5.79, p = .005). However, trait aggression (F [1, 57] = 1.05, p = .31), exposure to video game violence (F [1, 57] = 0.02, p = .89), and video-game ratings (F [1, 59] < 3.10, ps > .08) were not predictors of aggressive word completions.

Notably, in the data from Anderson et al. (2003) and Anderson et al. (2004) is the relevant p-values are all high-yet-significant. Specifically, each p-value is less than .05 and greater than .01. This would be extraordinarily rare if these were all tests of a non-null hypothesis (Simonsohn, Nelson, & Simmons, 2014). In Carnagey and Anderson (2005) there is a mix of significant and non-significant results. Collectively, these results do not strike us as unambiguously supportive of the validity of the AWCT.

However, implicit in these three studies is another form of self-justification for the AWCT. The AWCT is used to test the hypothesis that exposure to violent media (e.g., songs with violent lyrics, violent video games) increase aggressive cognitions. These studies argue that the AWCT is valid because AWCT scores increased after exposure to violent media. Thus, the AWCT is claimed to be a valid measure of aggressive cognitions because it supported these hypotheses AND these hypotheses were claimed to be supported because the AWCT is a valid measure of aggressive cognitions. The claim that violent media increases aggressive cognitions and the claim that the AWCT is a valid measure of aggressive cognitions do not stand independently. Rather, they each rest on the assumption the other is true. In short, this is a form of circular reasoning.

To throw another variable into the mix, let’s look at how the AWCT is administered in these three articles. Anderson et al. (2003) does not impose a time limit for administering the AWCT. Anderson et al. (2004) imposes a 3-minute time limit. Carnagey et al. (2005) imposes a 5-minute time limit. Thus, in these first 3 published studies using the AWCT, there are 3 slightly different procedures used to administer the task. The use of different procedures when administering the AWCT means that the evidence from these studies is not even directly cumulative.

So what?

We are not pointing to this pattern of citations to say “gotcha!” Rather, we are both aggression researchers who really want to ensure that our field is producing the best work possible. And, like a chef who ensures her knives are sharp or a mason who ensures his trowel is not bowed, we believe our best work is possible when we have confidence that our tools work well for the task at hand.

We believe that citing Anderson et al. (2003), Anderson et al. (2004), and Carnagey and Anderson (2005) is merely pointing out that this task has been used in previous publications. However, we believe these articles do not clearly demonstrate the AWCT is a valid measure of aggressive cognitions. In other words, these studies, in and of themselves, do not give us confidence in the AWCT as a useful tool for measuring aggressive cognitions.

Where do we go from here? As suggested by others (e.g., Zendle et al., 2018; see also Koopman et al., 2013), we strongly advocate for rigorous validation work on the AWCT. It may turn out that the AWCT is a valid measure of aggressive cognitions after all. If so, that would be great news. However, it may not. Either way, we need to know.

References

Anderson, C. A., Carnagey, N. L., & Eubanks, J. (2003). Exposure to violent media: The effects of songs with violent lyrics on aggressive thoughts and feelings. Journal of Personality and Social Psychology, 84, 960-971. DOI: 10.1037/0022-3514.84.5.960

Anderson, C. A., Carnagey, N. L., Flanagan, M., Benjamin, A. J., Jr., Eubanks, J., & Valentine, J. C. (2004). Violent video games: Specific effects of violent content on aggressive thoughts and behavior. Advances in Experimental Social Psychology, 36, 199-249. DOI:10.1016/S0065-2601(04)36004-1

Carnagey, N. L., & Anderson, C. A. (2005). The effects of reward and punishment in violent video games on aggressive affect, cognition, and behavior. Psychological Science, 16, 882-889. DOI: 10.1111/j.1467-9280.2005.01632.x

Koopman, J., Howe, M., Johnson, R. E., Tan, J. A., & Chang, C. (2013). A framework for developing word fragment completion tasks. Human Resource Management Review, 23, 242-253. DOI: 10.1016/j.hrmr.2012.12.005

Simonsohn, U., Nelson, L. D., & Simmons, J. P. (2014). p-Curve and effect size: Correcting for publication bias using only significant results. Perspectives on Psychological Science, 9, 666-681. DOI: 10.1177/1745691614553988

Zendle, D., Kudenko, D., & Cairns, P. (2018). Behavioural realism and the activation of aggressive concepts in violent video games. Entertainment Computing, 24, 21-29. DOI: 10.1016/j.entcom.2017.10.003

Monday, October 29, 2018

Harbingers of the replication crisis

I am going to post a series of tweets by James Heathers (who if you are not following on Twitter, you are really missing out) highlighting a series of passages dating back over half a century. They serve as warnings that, had they been heeded, would have left many of my peers in my corner of the sciences feeling a lot different about the soundness of the work we cite in our own research and teach to students:

























Let's not only heed those who tried to warn us over a half century ago, but heed those who are warning us now. I will not be around in a half century, but I will be around long enough to suss out if we're going to be actually progressing as a science or if we are just going to run over the same old ground.