Showing posts sorted by date for query zhang. Sort by relevance Show all posts
Showing posts sorted by date for query zhang. Sort by relevance Show all posts

Tuesday, April 16, 2024

Postscript to the Preceding: Weapons Priming Effect and Potential Allegiance Effect

As I noted in my prior post, almost all of the experiments explicitly testing for a weapons priming effect (short term exposure to weapons increasing relative accessibility of aggressive cognition) involve the primary General Aggression Model theorists (Anderson and more recently Bushman) and/or their graduate students or associates. We are a very small group of individuals. The methods we use are strikingly similar, both in terms of independent variables and dependent variables. Over the years, we used very similar protocols when running our experiments. We used the same theoretical basis for our work. I can find one researcher citing any of our work independent of our clique who successfully replicated our findings (Korb, 2016), and she never published her Master's Thesis, as far as I am aware. With very few exceptions (e.g., Deuser, 1994), our experiments consistently found statistical significance. I often wonder if there were more independent efforts to replicate our basic findings, and if they were unsuccessful I would be curious to know what they thought happened, especially if they used protocols similar to the ones that we used. Otherwise, what we have is something of a niche area of inquiry that likely started and ended with just our cohort. I don't find that especially comforting.

Footnote: I am quite aware that Qian Zhang, who had a weapons priming effect paper retracted (Zhang et al, 2016), does still look at the weapons priming effect, and although his lab's findings on the surface are consistent with our own, I simply discard that work as I do not trust his reported descriptive and inferential statistics. Let's just say that GRIM and SPRITE tests tend to uncover mathematically impossible descriptive statistics in too many of his lab's findings. I shall leave it at that.

Sunday, March 31, 2024

The Weapons Priming Effect: A Brief History and a Word of Caution

After Carlson et al. (1990) published their meta-analysis, the weapons effect was considered an established phenomenon. Short-term exposure to weapons appeared to increase the level of aggressive behavioral responses in lab and field settings compared to short-term exposure to neutral stimuli. After 1990, there has been a dearth of research examining the effect of weapons on aggressive behavioral outcomes. Instead, there appeared to be a shift to examining the cognitive underpinnings of the weapons effect. Although there were a couple experiments in which participants were administered a Thematic Apperception Test (TAT) arguably as an attempt to assess if participants thought more aggressively when exposed to weapons (e.g., Frodi, 1975), it would not be until the 1990s until a small group of social psychologists would more explicitly assess if the mere exposure to weapons or weapon images primed aggressive cognition utilizing techniques pioneered by the Cognitive Revolution in psychology. 

By the mid-1990s, some aggression researchers were using schema or associative network theories as a means to understand the impact of various aggression-inducing stimuli on aggressive cognition, affect, and behavior. Craig Anderson, for example, was already developing a theory known at the time as the General Affective Aggression Model (GAAM), shortened to General Aggression Model (GAM; Anderson & Bushman, 2002) by the start of this century. In that model, individuals store aggression-related information in the form of cognitive schemas and behavioral scripts. Exposure to stimuli theoretically believed to be associated with aggression or violence would prime these schemas or scripts, leading to an increased accessibility of aggressive thoughts, and potentially leading to an increase in aggressive behavioral responses. One interpretation of the classic Berkowitz and LePage (1967) weapons effect experiment was that those participants in the control room containing rifles were primed to think more aggressively and hence, under high levels of provocation, respond more aggressively.

The first explicit effort to test for a weapons priming effect was in the dissertation of Deuser (1994), a student of Craig Anderson. The experiments in that particular dissertation did not demonstrate a weapons priming effect at all. It did not matter whether participants were exposed to weapons or neutral stimuli. There was no evidence to support the theory that weapons would prime aggressive thoughts. The first published evidence of a weapons priming effect was Anderson et al (1996), although the weapons priming effect was more secondary to the main purpose of the article, which was to establish the General Aggression Model as a theory and to test the effects of uncomfortable heat on aggressive cognition, affect, and attitudes. The weapons priming effect was tested on those participants who were not exposed to uncomfortable heat. Participants were exposed to either weapon or neutral object images and were given a Stroop test to assess accessibility of aggressive cognition. Participants primed with a gun had more aggressive thoughts than those primed with a neutral object. The effect was fairly small, which seems to be a theme with this line of research, but definitely noticeable. 

That finding by Anderson et al (1996) was promising. The next step was to examine if weapons truly semantically primed aggressive thoughts. Anderson et al (1998) conducted two experiments. This is where I and Bruce Bartholow (both graduate students at University of Missouri at the time) come in. I had already been exposed to the schema and script theories described in the article's introduction, which helped designing experiments considerably. I did a lot of the legwork to find a reasonably sensitive measure of accessibility of aggressive cognition, and eventually settled on a version of the pronunciation task, in which participants read the target word into a microphone, and the time it takes between onset of stimulus and when participants' voices are picked up on the microphone is measured in milliseconds. For our purposes, participants demonstrated relative accessibility of aggressive thoughts if they reacted faster to aggressive target words than non-aggressive target words. The only difference between our two experiments were the stimuli. In Experiment 1, the prime stimuli were weapon words versus animal words. In Experiment 2, the prime stimuli were a mix of weapon images or a mix of neutral images. We had participants go through a few blocks of trials in which participants would first see the prime and then speak into the microphone when they saw the target word. So participants would first see a weapon (word or image) or neutral (word or image) concept and then saw an aggressive or non-aggressive target word, which they pronounced into the microphone as quickly and accurately as possible. We expected participants to show the fastest reaction times when weapon primes were paired with weapon target words. In each experiment, our expectations were confirmed. The effect size for Experiment 1 was in between small and medium, and in Experiment 2 closer to that of Anderson et al (1996). 

Since the publication of Anderson et al (1998), there have been several successful efforts to replicate and in some cases extend that finding. The dependent variables may differ (e.g., lexical decision task, aggressive word completion task or AWCT), and the experimental design may either be between-subjects or within-subjects, but the basic concept is the same. Bartholow et al (2005) in Experiment 2 started out as a replication and extension, and initial drafts of the manuscript included the analysis demonstrating the replication of Anderson et al (1998). That analysis was deleted in the published version. Thankfully, I kept the file with the analyses, although I don't have the original data file any more. There is a story behind the delay between when we ran our experiments and publication date, but that was more of an error by the editor, and can be chalked up to what life was like before electronic submissions of manuscripts. The effect was smaller than in any successful experiment up to that point, but still statistically significant. I think of the cognitive experiment by Lindsay and Anderson (2000) as a solid conceptual replication. Bartholow and Heinz (2006) would also offer a good faithful replication as a means of demonstrating that the concept of alcohol had a similar effect on relative accessibility of aggressive cognition. Subra et al (2010) would offer a conceptual replication of Bartholow and Heinz (2006). Bushman (2017) and Benjamin and Crosby (2019) also successfully replicated and extended the Anderson et al (1998) experiments. 

This seems like a fairly rosy picture. But here is where I start to have concerns. Every experiment I have mentioned thus far involves either Anderson or Bushman or their associates or former advisees. I have blogged before about experimenter allegiance effects before in the context of understanding the discrepancy between findings Berkowitz and his various colleagues and independent researchers who failed to replicate a weapons behavioral effect, including the Buss et al (1972) direct replication attempt. One of my concerns is that we may have a similar phenomenon with our weapons priming effect research, but there has been almost no work done to independently try to replicate our findings. The Anderson et al (1998) paper is one I am still proud of, and it is cited fairly frequently even today. But there is a problem. Almost anyone who is independent of our cohort of researchers uses that paper as a means of justifying their own experiments on phenomena that are often entirely divorced from our research. There may be an allegiance effect that we are missing. Most experiments outside of the Anderson-Bushman cohort that could arguably be coded as weapons priming experiments are often secondary to the main thrust of the published papers or use dependent variables that could defensibly be used as proxies of aggressive cognition, although reasonable skeptics would certainly have questions and might challenge any judgment that the measures involved are truly proxies of aggressive cognition. At least one experiment that was purported to be a conceptual replication of Anderson et al (1998) was later retracted due to the research and data analyses being fundamentally flawed, if not outright fraudulent (see Zhang et al, 2016, which I've written about before). I've seen one successful independent replication of Bartholow et al (2005) that used a somewhat defensible dependent variable (Korb, 2016). Regrettably, publication never came to fruition.

We are left with a cohort of weapons priming effect researchers who have almost exclusively used the GAM as a theoretical model (which is itself deserving of challenge), and may have produced experiments that are unique to our particular cohort. Independent researchers might take our protocols and find something entirely different. Such researchers may find different and arguably better ways to test the same hypotheses we tested. Unfortunately, there appears to be no way to know, unless or until enough aggression researchers come forward and report their own replication attempts either as preprints or as publications. If independent researchers find something different to what my group of researchers found, I think that would be interesting and worth understanding. I would want to know what was different. It is possible and likely probable that independent researchers would consider things we would not have considered? How that would impact the overall body of weapons priming research is very much unknown. It could we be the case that there was just something unique about any of us who studied the weapons priming effect who were either Anderson or Bushman and their various students/associates and that our findings can be safely ignored moving forward. All I know is that the findings from an experiment I conducted back at the start of my doctoral career (Anderson et al, 1998, Experiment 2), are in line with the overall effect size for aggressive thoughts that I reported in a meta-analysis (Benjamin et al, 2018).  With regard to that meta-analysis, I will gladly defend any coding decisions made in collaboration with my third author, even if we clearly differ in how to interpret our findings once we had ascertained that the analyses of effect sizes were sound. 

One final thought. Although I was trained via the GAM theoretical model, I am not wed to it. I think a good case could be made that the weapons priming effect is little more than a cognitive response equivalent to classical conditioning, and that outside of possibly cognitive responses, there isn't a whole lot left to write about. I've said that in some other contexts, so I'll say it here. There are theorists (including philosophers) who are trying to place the whole body of weapons effect research into different theoretical frameworks. Their work is worth examining. 

In the meantime, I am very concerned that my particular line of weapons priming effect research is little more than an experimental allegiance effect. I won't know more unless or until other independent researchers come forward. If their findings are consistent enough with what I and my cohort found, that's swell. We've at least established a cognitive response to what should be an aggressive-inducing stimulus. If not, aggression researchers like me need to go back to the drawing board. That is also okay. From my own perspective, I made my peace with this line of research several years ago. I started siding with the skeptics for a reason, and that was because the skeptics had the better argument on matters of theory and on the findings themselves. With regard to the weapons effect, or the weapons priming effect, I started out without a horse in that race. I will likely end my career without a horse in this race. Turns out there is nothing wrong with that.


Saturday, August 6, 2022

The struggle continues: Zheng and Zhang (2016) Pt. 4

Let's focus on Study 1 of Zheng and Zhang (2016). It should have been fairly simple, at least in terms of data reporting. However it happened, the authors honed into two video games that they thought were equivalent in terms of any confound aside from violent content, and merely needed to run a pilot study to demonstrate that they could back that up with solid evidence. It should have been a slam-dunk.

Not so fast.

The good news is that, unlike Study 2, the analyses of the age data are actually mathematically plausible. That's swell. I noticed that the authors had some Likert-style questions to rate the games on a variety of dimensions, which makes sense. The scaling was reported to be on a 1 to 5 scale, in which 1 meant very low and 5 meant very high for each dimension. My intention was to focus on Tables 1 and 2. If a 1 to 5 Likert scale was used for each of the items used to rate the games, there were some problems. One glaring problem is that there is no way that there could be means above 5. And yet, for Violent Content and Violent Images dimensions, the mean was definitely above 5 in each case. That does not compute. I have no idea what scaling was used on the questionnaires actually used. I can perhaps assume a 1 to 7 Likert scale. Certainly doing so would make some means and standard deviations that seemed mathematically impossible seem at least with in the realm of plausibility. But there is no way to know. We do not have the data. We do not have any of the materials and protocols. We have to take everything on faith. I had intended to have a set of images of SPRITE analyses on Table 1 and Table 2, but didn't see the point. 

Then we have the usual problem with degrees of freedom. With a 2x2 mixed ANOVA, with game type as a repeated measure and "gender" as a between-subjects factor, the degrees of freedom would not have deviated much from the sample size of 220. I think we can all agree with that. Degrees of freedom below 100 would be impossible. And yet the analyses reported do just that. It does not help much that Table 1 is mislabeled as t-test results. If we assumed paired sample t-tests, degrees of freedom for each item would have been 219. Again, the reported degrees of freedom do not compute.

What I can say with some certainty is that Zheng and Zhang (2016) should not be included in any meta-analysis addressing violent video games and aggression or media violence and aggression. My efforts to address some of these issues with the editorial staff never went very far. It's so funny how problems with a published paper lead editorial staff to go on vacation. I get it. I'd rather be out of town and away from email contact when someone emails (with evidence) concerns about a published paper. Unfortunately, if the data and analyses cannot be trusted, we have a problem. This is precisely the sort of paper that, once published, ends up included in meta-analyses. Meta-analysts who would rather exclude findings that are, at best, questionable will be pressured to include such papers anyway. How much that biases the overall findings is clearly a concern. And yet the attitude seems to be to let it go. The attitude is that the status quo is sufficient. One flawed study surely could not hurt that much? We simply don't know. The same lab persisted, with samples of over 3,000, to publish research relevant to media violence researchers. Several of those papers ended up retracted. Others probably should have been, but probably won't due to whatever political reasons one might imagine. 

All I can say is the truth is there. I've tried to lay it out. If someone wants to run with it and help make our science a bit better, I welcome you and your efforts.

Sunday, July 31, 2022

The NeverEnding Story: Zheng and Zhang (2016) Pt. 3

Whenever I have a few seconds of spare time and feel like torturing myself, I go back to reading a paper I have blogged about previously (see here and here). Each reading reveals more errors, and my remarks over the previous blog posts reflect that. Initially I thought Study 1 was probably okay or less problematic than Study 2. However, Study 1 is every bit as problematic as Study 2. I think I was so overwhelmed by the insane amount of errors in Study 2 that I had no energy left to devote to Study 1. And I do want to circle around to Study 1. But first, I want to add one more remark to Study 2.

With regard to Study 2, I focused on the very odd reporting of degrees of freedom (df) for each statistical analysis, given that the experiment had 240 participants. I showed that if we were to believe those df to be correct (hint: we shouldn't), there were several decision errors. And to top it off, the authors try to write off what appears to be a statistically significant 3-way interaction as non-significant. That would still be the case even if the appropriate df were reported. The so-called main effect of violent video games on reaction time to aggressive versus non-aggressive goal words was inadequate. As noted before, not only were the df undoubtedly wrong, but the analysis does not compare the difference in reaction times between the treatment and control conditions. I would have expected either a 2x2 ANOVA demonstrating the interaction or the authors to compute the differences (in milliseconds) between aggressive and non-aggressive goal words for both the treatment and control groups, and then to compute the appropriate one-way ANOVA or t-test. Anderson et al (1998) took this latter approach and were quite successful. At least the authors offered means for that main analysis. In subsequent analyses, the authors quickly dispense with reporting means at all. In no case do the authors report standard deviations. That's the capsule summary of my critique up to this point. Now to add the proverbial cherry on top: the one time that the authors do report the mean and standard deviation together was when reporting the age of the participants, and even then the authors manage to make a mess of things. 

Recall that the authors had a sample of 240 children ranging in age from 9 to 12 years for Study 2. The mean age for the participants was 11.66 with a standard deviation of 1.23. Since age can be treated as integer data, I used a post-peer-review tool called SPRITE to make sure that the mean age and standard deviation were mathematically possible. To do so, I entered the range of possible ages (as provided by the authors), the target mean and standard deviation, and the maximum number of distributions to generate. To my chagrin, I got an error message. Specifically, I was informed by SPRITE that the target standard deviation I had provided, based on what the authors reported, was too large. The largest mathematically plausible standard deviation was 1.17. Even something as elementary as the mean and standard deviation of participants' age gets messed up. You can try SPRITE for yourself and determine if what I am finding is correct. My guess is you will. Below is the result I obtained. I prefer to show my work.





So Study 2 is not to be trusted at all. What about Study 1? It's a mess for its own reasons. I'll circle back to that in a future post.

Friday, July 29, 2022

A Blast From the Past: Retractions and Meta-Analysis Edition

I stumbled across this article, Media and aggression research retracted under scrutiny, and found it to be an interesting short read. The article's author chronicles some recent retractions, and what had been another on-going investigation of several papers coauthored by Qian Zhang of Southwest University. I've written enough about his work over the last few years. I think referring to many of Zhang's papers having "been called into question" is a fair assessment. 

Part of the story chronicles Samuel West, who included one of Zhang's papers in his meta-analysis at the request of a reviewer. His meta-analysis would undergo another round of peer review around the time he learned of that particular Zhang paper being under investigation at the same journal. Ouch. West certainly has legitimate concerns about including a potentially dodgy finding in his meta-analysis. In this case, the paper by Zhang and colleagues was not retracted, but I am sure West has his misgivings about including the paper in his database in the first place. I can certainly empathize. My most recent published meta-analysis included one of Zhang's papers that would eventually get retracted early this year. That said, there are plenty of papers generated from Zhang's lab with obvious problems, or, in the case of his more recent work, have problems that are more cleverly hidden. I agree with Amy Orben that the fact that problematic studies continue to remain in journals and meta-analyses is "a major problem" when we think about how politicized media violence research is. Requiring archiving of data, data analyses, and research protocols probably helps to the extent that it is required - at least anything that might be incorrect or fraudulent can more easily be sniffed out. Otherwise, one can only hope for sleuths with enough time on their hands and no concerns for career repercussions for blowing the whistle on published papers that should have never seen the light of day. Good luck with that.

I do take issue with Zhang's characterization of Hilgard as someone who is "just trying to make his name based just on claiming that everyone else does bad research." I get that Zhang is a bit sore about the retractions, and Hilgard was the person who contacted Zhang and a plethora of journal editors regarding the papers in question. That said, there was plenty of chatter about Zhang's work in 2018 and onward, and there were probably several of us who just wanted to know that we hadn't gone insane, and that the obvious data errors, including degrees of freedom that were inaccurate, means and standard deviations that were mathematically impossible, and tables that made no sense really were what we thought they were. Hilgard was far and away better connected to the sphere of media violence research as an active researcher himself, and had the data analytic know-how and the connections that come with being at a R-1 university to do what needed to be done. Aside from that, Hilgard made plenty of positive contributions to the methodology side of psychological science, and from interacting with him online and in person over the years, I'll simply say he's a good person to know. 

I think this article is somewhat helpful in pointing out that even those who believe there is a link between violent content in media (such as video games) and aggression can view Zhang's work and see it for what it is, and express an appropriate level of skepticism. At the end of the day, one can take a philosophical perspective that there is "no one right way to look at the data" and that's all well and good. But at the end of the day, if the analyses show decision errors, and the means and standard deviations forming the basis for those analyses are simply mathematically impossible, the only reasonable conclusion that can be made is that the data and analyses in their present form cannot be accepted as valid. 

The only bone I really have to pick is that the author characterizes the body of media violence research as asking the question of whether or not "violent entertainment causes violence". Although I am aware that there are researchers in this area of inquiry who would draw that conclusion, there are plenty of other investigators who view what we can learn based on our available methods much more cautiously (a lot of aggression is mild, after all). There are also plenty of skeptics who doubt that there is any link between media violence and even the mild forms of aggression that we can measure. As far as I am aware, there is no link between exposure to violent content in mass media and violent behavior in everyday life. All that said, this is a useful article that captures a series of events that I know quite intimately. 

Suddenly, I am in the mood for some cartoon violence. I think I'll watch some early episodes of Rick and Morty. Goodnight.

Friday, February 4, 2022

A long-overdue retraction

After sounding the alarm bells several years ago, a paper that I had failed to get retracted (the editor of PAID at the time offered a superficially "better" Corrigendum in 2019 instead) is now officially retracted. Dr. Joe Hilgard really put the work in to make it happen. Here is his story:

The saga of this weapons priming article is over. There are plenty of articles remaining that have yet to be adequately scrutinized.

Monday, October 4, 2021

A grim day for another weapons effect paper

Sometimes a specific lab becomes the gift that keeps on giving. If the work is good, we are the better for it. If the work is questionable, our science becomes less trustworthy not only to the public, but to those of us who serve as educators and fellow researchers. As is true in other facets of life, there are gifts we would really rather return. 

Which brings us to a certain researcher from Southwest University: Qian Zhang. In spite of several recent retractions, I have to give the man credit. He remains prolific. A recent paper was recently uploaded on a preprint server, on a variation of weapons effect research that is quite well known to me. The author was even kind enough to upload the data and the analyses at osf.io, which is to be commended. 

That said, there are clearly some concerns with this paper. I will only discuss a few in this post. My hope is that others who are far more facile at error detection and have enough fluency in Mandarin can pick up where I will likely leave off. The basic premise of the paper is to examine if playing with weapon toys will lead children to show more accessibility of aggressive cognition (or think more aggressively, if that is easier on the eyes and tongue) and show higher levels of aggression on the Competitive Reaction Time Task (CRTT). As an aside, I seem to have some difficulty with the acronym, CRTT, and often misspell it when I tweet about the task. But I digress. These sorts of experiments have been run primarily in North America (specifically, the US) and Europe, and not so much outside of those limited geographical regions. Research of this sort outside of the US and Europe could be potentially useful if done well. Usually, experiments of this sort are done to examine only behavioral outcomes (Mendoza's 1972 dissertation is arguably an exception, if we code the variation of a TAT as a cognitive measure), so the idea of also examining cognitive outcomes could be potentially beneficial. 

As I read through the paper, I noticed that there were 104 participants in total. The author contends that he used an a priori power test to determine sample size at 95% power, using G-Power 3.1. That caught my attention. I dusted off my meta-analysis (Benjamin et al., 2018) and looked at effect size estimates for various distributions that we were interested in examining at the time. One of those distributions specifically included studies in which toy weapons were used as primes. The naive effect size is not exactly overwhelming: d = 0.32. That is arguably a generous estimate, once we include various techniques for measuring the impact of publication bias, and a good-faith argument can easily be made that the true effect size for this particular type of prime is close to negligible. But let's ignore that detail for a minute. Let's pretend we exist in a universe in which the naive effect size of d = 0.32 is correct. The authors argue that an N of 52 would suffice, but that their "sample size of (N=102 [sic])" was more than sufficient to meet 95% power. If you ever run any study in G-Power, you have to choose your analysis, enter the info required, and you are given a sample size estimate. One complication with G-Power is that it never directly allows us to enter an effect size for Cohen's d. It does give us Cohen's F. Computing Cohen's d from Cohen's F is quite easy: d = 2*F, and F = d/2. So, if I know that the effect size for my research question of interest is d = 0.32, I divide by 2 and can plug that into G-Power for Cohen's F, and then make sure I have my other info correct, including number of conditions, covariates, etc. When I do all that, based on the experiment as described, with a Cohen's F of 0.16, it becomes clear that the experiment would require a sample of at least 510 students. Now let's say that the author merely made a mistake and plugged in the number for Cohen's d by accident. The sample would still have to be about 129 in order to meet the requirements of 95% power, and really given the intention to randomly divide an equal number of males and females into treatment and control conditions, the author should shoot for 132 students. In order for the argument of 95% power to be met in this study, we'd have to assume a Cohen's d of approximately 1.00. There may be individual studies in the literature that would yield such a Cohen's d, but of the available sample of studies? Not so much. So, we have another low power experiment in the research literature. It's hardly the end of the world. 

What grabbed my attention was the research protocols described in the experiment. For the time being, I will take the author at his word that this was an experiment in which random assignment was involved (this author has once been flagged for failing to disclose that participants chose which treatment condition they were involved in, which was, shall we say, a wee bit embarrassing). The way the treatment and control conditions are described seems standard enough. What was odd was what happened after the play session ended. The children were first given a semantic classification task. I admit that I've had to do a double take on this, as some of the wording is a bit off. I am increasingly thinking that what the authors did was use a series of aggressive and neutral pictures and had children respond to them as quickly as they could. The author had made some mention of aggressive and neutral pictures also being used as cues, which threw me, because that would have seemed more like an experiment within an experiment. At minimum, there would have been needless contamination. Then the children participated in a CRTT where they set noise blasts at 70 to 100 db. Those controls were set from 0 (no noise) to 4 (100 db). The authors reported their means and standard deviations. I then initially looked at the means for treatment and control condition using GRIM, which is a nifty online tool for flagging errors. The results were, to say the least, initially looked grim. However, I was reminded that there is the issue of granularity that I might have overlooked. So, even though there is one scale, the trials each count as independent items. So, an N=26 for one cell is, with the 13 out of 25 trials that the author included in the data set (in which participants had an opportunity to send noise blasts after a loss), effectively an N=338. So I went and opened up the SPRITE test link and entered the same mean info, along with the minimum and maximum scale values (0 and 4, respectively), the target mean for each cell I was interested in, and SPRITE would report that each of the two cells measuring boys failed to arrive at a solution for at minimum the standard deviation. In each case, the standard deviation was reported by SPRITE to be too low. I can get reproductions of possible distributions for the other two cells. I then downloaded the data set to see what it looked like. Much of it is in Mandarin, but I can make some educated guesses about the data in each column. I turned my attention to the "ANCOVA" analysis. It actually looked like a MANOVA was run. Perhaps a MANCOVA (but as I am admittedly not literate in Mandarin, it's hard to really know without taking time I don't have yet to put some terms into Google Translate and sort that all out). That's a project for later in the month. I could see the overall mean for the aggressive behavioral outcome, as measured by the CRTT and entered it, its standard deviation, and overall N in SPRITE and noticed it also could generate some potential score distributions. Still, given the failure to generate some potential distributions in SPRITE, it's not a good day to have posted a preprint. At minimum, there is some sort of error in reporting, whatever the cause.

I do need to take some time to sit down, try to reproduce the analyses, etc. Of course, it goes without saying that successfully reproducing a data set that has been in some way fabricated is going to add no new information. That said, I am satisfied that the findings as presented for the CRTT analyses intended to establish that weapon toys could (at least most specifically for the male subsample) influence aggressive behavioral outcomes may also be potentially questionable. This is a paper that should not make it past peer review in its present form. 

Note that I have not yet run this through Statcheck, although in recent years Zhang's lab has become more savvy about avoiding obvious decision errors. I made an effort as of this writing to run the analyses as they appeared in the pdf, and the report came back with nothing to be analyzed. I will likely have to enter the analyses by hand on a word document and then reupload at a later date. 

Please also note that the author appears not to have counterbalanced the SCT and CRTT measures to control for order effects. That strikes me as odd. The very superficial discussion about debriefing left me with a few questions as well. 

Note: Updated to reflect some more refined analyses. Any initial mistakes with GRIM are my own. I am on solid ground with the SPRITE runs, and I think my own concerns about the lack of statistical power, failure to counterbalance, etc. are on solid ground.

Saturday, July 10, 2021

Weapons Effect Theory

A little while ago, I made mention that I had noticed the weapons effect, which I had always considered to be a phenomenon, referred to as a theory. In a way I found it amusing. In another way, I think the argument in favor of a weapons effect theory does have some merit. A good theoretical model would at minimum offer an explanation regarding a phenomenon and generate testable hypotheses. In the case of the weapons effect, it would be a relatively narrow theory. Then again, so too was frustration-aggression theory (itself an outgrowth of what was merely a hypothesis). 

We know the origins of what we could call weapons effect theory. We look no further than Berkowitz and LePage (1967). As the details of that initial experiment are detailed elsewhere, I will simply state that Berkowitz and LePage (1967) appeared to demonstrate that under conditions of high provocation, individuals experiencing short-term exposure to weapons showed higher levels of aggression (measured in number of electric shocks given) than those who had not been exposed to weapons. It goes without saying that the claim was highly controversial at the time, and that there were critics who could not replicate the original finding. That story has been told many times (including by me - see Benjamin 2019 or Benjamin, 2021), and bears no repeating here. What probably matters most is that a meta-analysis by Carlson, Marcus-Newhall, and Miller (1990) was supposed to have settled the matter. Short-term exposure to weapons under conditions of high provocation or frustration seemed to lead to a noticeably higher level of aggression than any other condition. Armed with Fail-Safe N as a means of assessing publication bias, Carlson et al (1990) concluded that the case was effectively open and shut. The weapons effect was viable, and it was time to move on. After that, social psychologists and some fellow travelers looked toward underlying processes responsible for this purported effect. That's where theory comes in.

Although the Anderson, Benjamin, and Bartholow (1998) paper referenced Anderson's then General Affective Aggression Model (which would be later abbreviated to General Aggression Model or GAM), I think it is safe to say that what we we actually did was to articulate a distinctive weapons effect theoretical model. Among social cognition models, it is a potentially "warm" theory in the sense that anger and arousal are considered potential antecedents. However, anger (affect) and arousal have never been adequately tested. Rather, testing of the model primarily focuses on short-term exposure to weapons priming of aggressive cognitions - think of these as behavioral scripts and schemas that include all of our semantic and episodic memories and concepts of aggression and violence as well as procedural memories of how to behave aggressively or violently. These memories may be implicit or explicit. Once aggressive cognitions have been primed, primary and secondary threat appraisals are primed, increasing the likelihood that an individual will be biased to perceive stimuli as more threatening than they might have otherwise, along with appraisals of how to best respond. Depending at what happens at the level of appraisal, an aggressive behavioral response might be the end result. Although primarily focused on the situational antecedents, the model keeps the door open to individual differences that might serve as antecedents (including personality traits and life experience). See the figure below. Note that technically this figure is the property of Sage Publications (from Anderson et al,, 1998), and if I am asked to take it down, I will do so:



It's a simple theory, really. One sees a weapon, which facilitates an increase in accessibility of aggressive cognitions, setting up primary and secondary threat appraisals, culminating in increase of aggression. The potential for weapons to prime anger and increase physiological arousal exist as well. It is a model that explains a body of results on a phenomenon, and offers some potential hypothesis tests. So far, so good. So, how well does the weapons effect theory hold up? Depending on whom you read, the weapons effect theory is either sufficiently established that we what we really need to do is to further explore interactions of person variables and short-term exposure to weapons (an endeavor that has barely been undertaken, and then only in a very scattershot fashion), or the body of research suggests the theory is enough of a nothingburger as to be swept into the dustbin of history.

When I finally published my meta-analysis (Benjamin, Kepes, & Bushman, 2018), I think any astute reader would hone in on Table 2 and realize that depending on how how publication bias is assessed, that there is nothing to be concerned about (if one believes random-effects trim-and-fill analyses) or quite serious (e.g., PET-PEESE). Most concerning are studies examining aggressive behavioral outcomes. The effect sizes are arguably negligible. Even when we look at the intervening variables in the model that are the underlying processes responsible for the presumed relationship between short-term weapon exposure and aggression (accessibility to aggressive cognitions and hostile appraisals) we have to keep in mind that the effects for these outcome variables are often small. Establishing accessibility of aggressive cognition is difficult, and numerous methods of measuring accessibility of aggressive cognitions have been utilized with varying degrees of success. Although much of the earlier cognitive priming literature for the weapons effect relied on either reaction times to aggressive versus non-aggressive words in lexical decision tasks or pronunciation tasks, more contemporary studies appear to rely on variations of a word completion task developed by Anderson - an instrument whose validity has been recently questioned. I wonder how many unpublished studies slipped through the cracks. Research on mostly primary threat appraisal has been more of a success story. Much of that work seems to build on research comparing phylogenitic and ontogenetic threats, with weapons being an ontogenetic threat. When individuals are shown arrays of objects with guns or knives embedded, studies appear to find evidence that individuals respond more rapidly to those arrays that have nothing but neutral objects. Effect sizes are small-to-moderate. One must also consider the possibility that arrays including unexpected objects could be just as effective in decreasing reaction times. So, although the pattern of findings looks promising, it's probably far from settled. But ultimately, for a cognitively based theory of the weapons effect to work, there has to be some establishment that aggressive behavioral outcomes are consistently positive. So far, that has not been the case. However aggressive behavior has been operationally defined - number of electric shocks, shock/noise blast levels, amount of hot sauce doled out to a presumed victim, point subtraction, etc., the results have been inconsistent. Some experiments appear successful, but many others do not replicate - either directly or conceptually - the original finding. Furthermore, it is not entirely clear that we are measuring aggression with these operational definitions, nor can we necessarily include intention to harm from the body of research thus far. Furthermore, there has been a tendency for those who do find positive effects to oversell their findings, tying their analyses to not merely the mild forms of aggression that we might be measuring, but to tie that work to acts of violence such as shootings. The theoretical model is not one that was designed to address violence per se, which means that even if we could hone in on consistently reliable findings, we can only speak to a narrow range of possible aggressive behaviors in everyday life, and even then with a good deal of caution and humility. All that being said, we have  a social cognitive model that states that short term exposure to weapons can trigger an aggressive behavioral response, to the extent that aggressive cognitions and hostile appraisals are successfully primed. Without a solid body of evidence pointing to an increase in aggressive behavior in this body of research, the theory falls apart. Unless or until the behavioral outcome piece of the theory is settled, it is a weak theory at best. If research surfaces that debunks any priming effect of weapons on aggressive cognitions, then the theory goes from weak to effectively moot. At that point, why even discuss the matter further?


Left unanswered in the theoretical model is the role of arousal. My impression is that early on, arousal was looked at as a nuisance variable to be measured and ruled out. I know of one explicitly reported arousal study, and it was a pilot study used to select stimulus materials.  As a "warm" theory, I've been a bit taken aback at the lack of interest among those best positioned to examine arousal and affect to actually take the time to do so and report their findings. Nor has the moderating role of individual differences been adequately explored. Aside from some one-offs, very little is known about the role of personality or life events as a moderator of the relationship between short-term exposure to weapons and aggressive behavioral outcomes. 

Research from the last decade has been discouraging. In the last two or three years, we've seen published some ecological valid behavioral work that was either adequately sampled, but showed a small effect size, or what on the surface appeared to be a robust effect, but in which the sample small enough that the statistical test was underpowered. Let's just say that historically experiments testing hypotheses derived from this theory rarely have samples of 15 or more in each cell, and even a sample of 15 per cell is probably inadequate. More traditional experimental research (i.e, in the lab) in recent years appears to suggest the behavioral effect is minimal at best. For example Guo, Egan, and Zhang (2016) found no main effect of weapons on aggressive behavior, and instead used a subsample of individuals who scored as high in external locus of control in order to craft a narrative for their findings. That's just the published research. I am aware of unpublished behavioral experiments (based on personal communication) that have found either a null effect, or even a suppression effect. That should give any of us who either have researched the weapons effect theory or who utilize this theory as part of our pedagogy pause.

Bottom line? As a theoretical model, I am very uncertain that the weapons effect theory is on solid ground. If anything, I am likely to agree with those who would argue that it is not on solid ground at all, and that it is a theoretical model worth abandoning. There appear to be small to moderate effects when it comes to weapons priming aggressive cognitions and hostile threat appraisal. The effect on aggressive behavior appears potentially negligible. I say that as someone whose professional identity was in some significant sense tied to this particular theory. I also say this as someone who has, in the past taught history of psychology to undergraduates, and who has a keen interest in my area's history. The conditions that made a link between short-term exposure to weapons and the very mundane aggression we can observe in the lab are ones in which there was both an increase in violent media consumption (in which weapons were ubiquitous) and an increase in real life violence (something Berkowitz goes into in a 1968 paper) seemed plausible. Since that time, the concept of media violence and real life violence has been effectively debunked. Whether or not a weapons effect theory holds up in the sense that, say frustration-aggression appears to hold up is questionable at best. I think a registered replication report of the original Berkowitz and LePage (1967) would be wise, assuming it were ethically and logistically doable, if for no other reason than to settle the matter once and for all.

Friday, January 29, 2021

Research Confidential: How Self-Correcting is Science?

The question is arguably rhetorical. Science in and of itself is not self-correcting. It takes living, breathing human beings to notice something is wrong, take the time and make the effort to report what is wrong to relevant stakeholders (e.g., journal editors, relevant university adminstrations, etc.), and then have good reason to believe that the relevant stakeholders will show due diligence, correct or retract flawed papers as needed, and otherwise hold those responsible for the flaws, whether due to sheer incompetence or fraud, accountable. In an ideal world, that is how it would work. In this world, it's considerably more complicated, and often more than a bit disheartening.

If you are a regular reader of this blog, you are quite aware of our favorite media violence researcher who is notorious for some of the worst papers published in that particular niche area of psychology - Qian Zhang of Southwest University. I have documented, over the last couple years or so, some of the most insane tables with means, standard deviations, and test statistics that simply are impossible to interpret. I have reported test statistics that, based on the degrees of freedom reported, would have to be incorrect. I have reported discrepancies between degrees of freedom for test statistics and the sample size reported. I have documented evidence of potential plagiarism and self-plagiarism - the latter due to the tendency for Zhang to rely heavily on copying and pasting from one paper to another. I have also found some amusing typos that resulted from Zhang's tendency to copy and paste tables from paper to paper. I've tagged Zhang's work as I have documented here (for your convenience) and on PubPeer under a pseudonym. 

Dr. Joe Hilgard has gone considerable further than have I. He's blogged about his own experiences in documenting problems with Zhang's work in great detail, and the efforts he's made to contact journal editors along with officials at Zhang's university, offering painstaking evidence of the problems he has discovered. You can read about Joe Hilgard's efforts, and the decidedly mixed and disappointing outcome of his efforts here. You should really take to the time to read Hilgard's post as it is thorough and damning. The short version? Some journal editors responded rather well, and in one case very quickly to retract two papers that were clearly unsound. Other journal editors have either stonewalled or ignored Hilgard's concerns. Zhang's university cleared him of wrongdoing, chalking it all up to Zhang being "deficient in statistical knowledge and research methods." So in other words, the university writes it off as "the guy's merely an idiot, but hey, let's just give him a remedial stats course and call it even." I agree with Hilgard that the university's failure to take action is not that surprising, as universities seem to be in the business of taking care of their own, especially if the researcher in question might be bringing in grants or other forms of prestige. So the guy maybe fudges some numbers and has no idea what random assignment means. There's nothing to see here. Move along.

My take on the matter is that the most charitable view that one could take based on the body of Qian Zhang's work is that this is a researcher who is grossly incompetent, but that a more probably defensible case can be made that his activities are on some level fraudulent. I am more inclined to the latter less charitable view. I've seen too much. Regardless, this is research that should have never made it past peer review. I agree with Hilgard that this body of research is very problematic given that as long as it remains published, it will distort our understanding of what is actually happening with stimuli such as video games that contain violent content on outcome variables such as aggressive behavior or cognition. Meta-analyses are especially vulnerable given that some of the reported findings by Zhang rely on large samples. Those results could artificially inflate effect sizes, leading meta-analysts and those consuming meta-analyses to believe that an overall effect is stronger than it actually is. 

This is one of the dark alleys I mentioned a few years ago. And given what Hilgard has experienced and what I've experienced in my own way, it's one that few leave with any sense of hope for the state of this particular are of psychological inquiry. If blatantly problematic papers, ones where the problems are so obvious that a beginning methods student could discover them, cannot be retracted within a short window of time, what is going on with work in which potentially fraudulent data analyses are more cleverly presented? What else is out there that cannot be trusted? That is something that should cause us all to lose some sleep.

One final thought for anyone thinking of collaborating with Zhang: don't. If you absolutely cannot help yourself, insist on seeing the data before agreeing to be part of that particular project. I'd say that is a safe practice regardless of the situation. If I take on a statistician for a project, or someone who is at least better versed in a particular statistical method than I am, I insist on sending the data set or database, and I expect that the statistician on the project will double check my work and ask difficult questions as needed. That can save a lot of grief, assuming that the statistician involved is actually looking at what is being sent. One of the tragedies for some of Zhang's coauthors is that they've never had access to the data sets to which they lent their names and reputations, nor were they apparently allowed access. That is not how we do science, folks.

In the meantime, more papers are in the pipeline to be published by this particular author, and it will become more of a struggle to keep up with the dross that is likely to be found in any of those papers. Again, that is something that should cause us all to lose some sleep.

Monday, December 7, 2020

And yet another one

Zhang lab has yet another article published online. It can be found here:

https://doi.org/10.1016/j.jecp.2020.105005

On initial inspection, it appears to have a similar format, and problems that another recent paper from the same lab has. I suspect this story will continue to develop.

Thursday, October 8, 2020

Postscript to the preceding: This isn't the first odd mistake for Zhang

 Previously, I noted that the latest Zhang et al. (2020) paper had at least one serious error: that instead of computing a difference between reaction time for aggressive (weapon) images and neutral images, the authors used simply the reaction times to the weapon images as the DV. Hence, we as the readers are left with a misleading set of analyses and a potentially misleading narrative. Fortunately, the authors had already shared their data, which made detecting the error fairly easy. Why the reaction times for the neutral images and then the difference scores (which would have been the real DV) didn't have their own column is only something that the authors can answer.

Oftentimes, with this lab (as is probably the case with others), it is often difficult to glean whether or not variables are entered and computed correctly based on the information appearing in a published paper. Whether those omissions are a bit of sleight of hand or simple human error or misunderstanding is often difficult to deduce. However, sometimes authors make it easy for the readers to see for themselves that the authors have goofed. I have found some rather odd analyses in which IVs were not quite analyzed correctly as well as DVs.

One of my favorite papers published by the Zhang lab, just for the sheer madness it contained, was the on published in Personality and Individual Differences nearly five years ago. That was the first, and I think only, effort these authors made to replicate and extend research on the weapons priming effect (itself a fairly controversial topic). The DV situation appears okay in the initial analysis under section 4.1. However, where things fall apart (aside from a grossly undersized df, given sample size) was that the authors only examined the difference in reaction times between aggressive and neutral words under the weapon prime condition, while completely ignoring the neutral prime condition. The authors eventually did correct the df for that section in a pretty massive corrigendum. However, they never did address that they had done the wrong analysis in order to establish a weapons priming effect. They really should have read more carefully Anderson et al. (1998) in order to do so. The authors needed to establish that the difference between rts in the treatment and control conditions were larger, and in the predicted direction, when participants saw weapons than when they were presented with neutral images. Also left unanswered was the nagging question of the three-way interaction effect that was a duplicate of another three-way interaction effect in another paper authored by this same research team. I got the impression that the current editor in chief at Personality and Individual Differences was not much in the mood for dealing with this mess to begin with, and that any superficial corrections were extracted from Zhang et al. (2016) was probably a minor miracle. In theory, since the authors changed a single digit in the F-test for the three-way interaction, perhaps the point is now moot. I am still concerned that a certain amount of self-plagiarism happened, but the editor-in-chief chose to let it go. As was the case with the most recent article in question, the Zhang lab had enlisted an established American aggression researcher, Phillip Rodkin. Rodkin's wheelhouse was more in the area of bullying, and not so much media violence, so this seemed like an odd choice for a collaborator for a media violence paper. I honestly don't know how much access Rodkin had to the original data, nor could I comment on whether he would have known what to look for when checking out the analyses. He had already been deceased for a while when this paper was published. Hence, we will likely never know.

The Zhang et al. (2016) paper shared something strikingly in common with a paper in which Zhang was second author, and Rodkin also was a collaborator. There was a three-way interaction that was deemed nonsignificant in each paper, although according to a Statcheck analysis, the three-way interaction would have to have been statistically significant based on what was originally reported in each paper. Publication of duplicate analyses is presumably serious business, but apparently the powers that be can overlook such matters. Perhaps the corrigendum on the Zhang et al (2016) paper makes the point moot, as I noted earlier. The erratum in the other paper entirely ignores the pesky issue of that three-way interaction effect. 

As I have probably said too many times, I find this state of affairs to be very disappointing. As someone who still finds media violence research interesting (although definitely from the standpoint of a skeptic), I treasure efforts by researchers who study non-WEIRD populations. As an educator and researcher who is very eager to decolonize my particular areas of expertise, I would ordinarily welcome work coming out of China. Unfortunately, the work from this lab is so chock full of errors that it is best left uncited. Hold out for the real thing. Hold out for competently and ethically conducted work. 

References:

https://doi.org/10.1016/j.paid.2015.09.017 

https://doi.org/10.1016/j.paid.2018.12.010 

http://dx.doi.org/10.4236/ojmp.2016.52005

http://dx.doi.org/10.4236/ojmp.2019.83005

Wednesday, October 7, 2020

The Zhang Lab rides again

If you've read this blog long enough, you're familiar with the work of Qian Zhang of Southwest University in China. You are already well aware that there are some serious problems with many of the papers he has co-authored (either as a first author or a more secondary co-author) over the years. His more recent papers have been on the surface of higher quality, but it sometimes doesn't take much to realize that there are still substantial problems. Bottom line is that if you see his name mentioned here, it's not good news.

Case in point: Dr. Zhang has a new paper out that purports to examine the link between viewing prosocial cartoons and a reduction in aggressive cognition and behavior. I was alerted to this paper by Joe Hilgard. On the surface, a simple Statcheck run looked good. Initially I lamented the lack of tangible data to reproduce the analyses. Dr. Hilgard pointed me to where the data were stored (which kudos to this lab for doing so). I ported the dataset into my current version of jamovi and successfully reproduced the analyses reported in the paper. So far so good. Then I had that sinking realization something was still wrong. The data set only contained data for reaction time data for weapon images (which the authors use for the DV in the paper and analyses - a fact Dr. Hilgard had already arrived at before I did my work here). However, the authors should also had data on reaction times for neutral images that were not included in the data set. The appropriate DV would have been a difference score between reaction times for aggressive images (in this case, weapons) and reaction times neutral images. That difference score would be the proper measure of accessibility of aggressive cognitions. 

As of this writing, the last author on the paper had been contacted, and I trust this last author to do the right thing here. At bare minimum, a reanalysis needs to be conducted in order to ascertain that prosocial cartoons really did lead to a decrease in the relative accessibility of aggressive cognition. As of now, the paper cannot adequately address that claim. There is this funny gray area between what we consider published and in press. The paper has already been accepted, and some version of it has been made available online. My hope is that the last author, an American researcher with a solid reputation, is able to get the matter resolved satisfactorily, however that turns out. Maybe a simple correction suffices. It is possible that a properly calculated cognitive DV yields the same basic findings as the incorrect cognitive DV. If it does not, then many of the conclusions of the paper may need to be rethought and rewritten. If so, the topic is of enough theoretical and practical interest that perhaps a sympathetic editor and publisher will still be okay with a corrigendum, regardless of how the ultimate findings flush out. If a retraction is necessary at this stage, it would be far less painful than after it is already officially in print. This is a matter of making sure that those of us who might still be tempted to conduct meta-analyses in this broad area of media violence have the correct findings when estimating effect sizes, that those who might be using this literature to advocate for policy changes have the right information before coming across as grossly uninformed. For the good of the order, I hope this matter is taken care of quickly. In the meantime, I'd warn against citing this particular paper unless and until at least some sort of correction has been published. 

Reference: https://doi.org/10.1016/j.childyouth.2020.105498

Thursday, July 9, 2020

Update - what's happened with Zhang Lab papers?

Short answer is that not a lot has happened since the end of last year. More to the point: nothing seems to have happened. I have seen no new English-language publications from the lab. Maybe some specifically Chinese publications have emerged. As of now, I am unaware of them. As of now, there are two retractions (both of which involve reputable scholars, and for whom I can only offer my sympathies), a Corrigendum (which turned out to be only a partial solution - the article needs to be retracted), and a number of errata in what are potentially predatory journals, that are themselves in need of errata. There are a couple journals that have yet to issue so much as a message of concern, despite some glaring errors. If nothing else, the web presence of Qian Zhang at Southwest University in Chongqing has changed considerably over the last year or so. At one point, Zhang had several photos of himself and with eminent US Psychologists from Illinois, along with some statement about SPSS expertise and an enticement to potential grad students to work in his lab. All of that is gone. There is some description of past work, including recent. That's it. Maybe that is progress of a sort. I have no idea of what the CCP has in mind for this particular researcher, nor any particular concern either way. My main concern is that I and my peers can compute accurate effect size estimates, reproduce the findings, and replicate the work. My impression, based on doing a StatCheck scan on a Chinese-language article just prior to Qian Zhang joining the lab at Southwest University, is that the pattern of errors was already in place. This is someone who apparently adapted to a lab culture that was itself in need of improvement. In a toxic academic environment (of which many of us are all too familiar) the convenient way of getting along is the path of least resistance. I wish it were not that way. My guess is that we are looking at a tragedy of errors - one in which there are no villains, just people who made a lot of regrettable choices for the same reason anyone might make regrettable choices in a late capitalist economy. If we get this right in our corner of the sciences, this lab's body of work will be a cautionary tale about the role incentive structures play in terms of career advancement. I suspect there is also a cautionary tale about how eminent psychologists grease the path for success of ambitious researchers, regardless their actual talent and research practices. There is a cautionary tale of coauthors having thorough access to data and codebooks. There is a cautionary tale about editors and peer reviewers having the tools at their disposal to to their jobs as well as possible.There are no happy endings for this particular case. For those of us who really do treasure samples of non-WEIRD populations, I advocate only for making sure that the protocols and data analyses are above board. Else, we get a situation that is truly a mess, and one in which any scholar wishing to extract effect sizes for meta-analyses in the broad area of media violence will be left flustered.

Friday, December 6, 2019

Now about those Youth and Society retractions involving Qian Zhang

Hopefully you have had a moment to digest the recent article about the retraction of two of Qian Zhang's papers in Retraction Watch. I began tweeting about one of the articles in question around late September, 2018. You can follow this link to my tweet storm for the now retracted Zhang, Espelage, and Zhang (2018) paper. Under a pseudonym, I initially documented some concerns about both papers in PubPeer: Zhang, Espelage, and Zhang (2018) and Zhang, Espelage, and Rost (2018).

Really, what I did was to upload the papers into the online version of Statcheck, and flag any decision inconsistencies I noticed. I also tried to be mindful of any other oddities that seemed to stick out at that time. I might make note of df that seemed unusual given the reported sample size, for example, or problems with tables presuming to report means and standard deviations. By the time I would have looked at these two papers, I suspect that I was already concerned that papers from Zhang's lab showed a pervasive pattern of errors. Sadly, these two were no different.

With regard to Zhang, Espelage, and Zhang (2018), a Statcheck scan showed three decision errors. In this case, these were errors where the authors reported findings as statistically significant when they were not - given the test statistic value, the degrees of freedom, and the level of significance the authors tried to report.

The first decision inconsistency has to do with an assertion that playing violent video games increased accessibility of aggressive thoughts. The authors initially reported the effect as F(1, 51) = 2.87, p < .05. The actual p-value would have been .09634, according to Statcheck. In other words, there is no main effect for violent content of video games in this sample. Nor was a video game type by gender interaction found: F(1, 65) = 3.58, p < .01 actual p-value: p = 0.06293. Finally, there is no game type by age interaction: F(1, 64) = 3.64, p < .05 actual p-value: p = 0.06090. Stranger still, the sample was approximately 3000 students. Why were the denominator degrees of freedom so small for these reported test statistics? Something did not add up. Table 1 from Zhang, Espelage, and Zhang (2018) was also completely impossible to interpret - an issue I have highlighted in other papers that have been published from his lab:

A few months later, a correction would be published in which the authors would purport to correct a number of errors found on several of the pages of the original article, as well as Table 1. That was wonderful insofar as it went. However, there was a new oddity. The authors purported to only use 500 of the 3000 participants in order to have a "true experiment" - which was one of the more interesting uses of that term I have read over the course of my career. And as Joe Hilgard has aptly noticed, problems with the descriptive statistics continued to be pervasive - implausible and impossible cell means and marginal means and standard deviations, for example.

With regard to the Zhang, Espelage, and Rost (2018) article, my initial flag was simply for some decision errors in Study 1, in which the authors were attempting to establish that their stimulus materials were equivalent across a number of variables except for the level of violent content, and consistent across subsamples, such as sex of participant (male/female). Given the difficulty that exists in obtaining, say, film clips that are sufficiently equivalent except for level of violence, due diligence in using materials that are as equivalent as possible, except for the IV or IVs is to be admired. Unfortunately, There were several decision errors that I flagged after a Statcheck run.

As noted at the time, contra the authors' assertion, there was evidence that there were some rated differences between the violent film (Street Fighter) and the nonviolent film (Air Crisis) in terms of pleasantness - t(798) = 2.32, p > .05 actual p-value: p = 0.02059 - and fear - t(798) = 2.13, p > .05 actual p-value: p = 0.03348  To the extent that failure to control for those factors might have impacted subsequent analyses in Study 2 is of course debatable. It is clear that the authors cannot demonstrate, based on their reported analyses, that they had films that were equivalent on variables that they identified as important to hold constant with only violent content varying. The final decision inconsistency suggested that there was a sex difference in ratings of the variable fear, contrary to authors claim, t(798) = -2.14, p > .05 actual p-value: p = 0.03266. How much that impacted the experiment in study 2 was not something I thought I could assess, but I found it troubling and worth flagging. At minimum, the film clips were less equivalent than reported, and the subsamples were potentially reacting differently to these film clips than reported.

Although I did not comment on Study 2, Hilgard demonstrated that there was a consistency in the pattern of reported means in this paper were strikingly similar to the pattern of means reported in a couple of earlier papers in which Zhang was a lead or coauthor in 2013. That is troubling. If you have followed some of my coverage of Zhang's work on my blog, you are well aware that I have actually discovered at least one instance in which a reported test statistic was directly copied and pasted from one paper to another. Make of it what you will. Dr. Hilgard was able to eventually get a hold of the data that were to accompany a correction to that article, and as noted in the coverage in Retraction Watch, the data and analyses were fatally flawed.

I was noticing a pervasive pattern of errors in these papers, along with others I was reading by Zhang and colleagues at the time. These are the first papers on which Zhang is a lead or coauthor to be retracted. I am willing to bet that these will not be the last, given the evidence I have been sharing with you all here and on Twitter over the last year. I have already probably stated this repeatedly about these retractions - I am relieved. There is no joy to be had here. This has been a bad week for the authors involved. Also please note that I am taking great care here not to assign motive. I think that the evidence speaks for itself that the research was poorly conducted and poorly analyzed. That can happen for any of a number of reasons. I don't know any of the authors involved. I have some awareness of Dr. Espelage's work in bullying, but that is a bit outside my own specialty area. My impression of her work has always been favorable, and these retractions notwithstanding, I see no reason to change my impression of her work on bullying.

If I was sounding alarms in 2018 and onward, it is because Zhang had begun to increasingly enlist as collaborators well-regarded American and European researchers, and was beginning to publish in top-tier journals in various specialties within the Psychological Sciences, such as child and adolescent development and aggression. Given that I thought a reasonable case could be made that Zhang's reputation for well-conducted and analyzed research was far from ideal, I did not want to see otherwise reputable researchers put their careers on the line. My fears to a certain degree are now being realized.

Note that in the preparation of this post, I relied heavily on my tweets from Sept. 24, 2018 and a couple posts I published pseudonymously in PubPeer (see links above). And credit where it is due. I am glad I could get a conversation started about these papers (and others) by this lab. Joe Hilgard has clearly put a great deal of effort and talent into clearing the record since. Really we owe him a debt of gratitude. And also a debt of gratitude to those who have asked questions on Twitter, retweeted, and refused to let up on the pressure. Science is not self-correcting. It takes people who care to actively do the correcting.

For those visiting from Retraction Watch:

Retraction Watch posted an article about two retractions of articles in which Qian Zhang of Southwest University in China was the lead author. Since some of you might be interested in what I've documented about other published articles from Zhang's lab, your best bet is to either type Zhang in the search field for this blog. Or just follow this link, where I have done the work for you. I'll have more to say about these specific articles in a little bit. I think I documented some of my concerns on Twitter last year and pseudonymously on PubPeer. In the meantime, I am relieved to see two very flawed articles removed from the published record. Joe Hilgard deserves a tremendous amount of credit for his work reanalyzing some data he was able to obtain from the lab (and his meticulous documentation of the flaws in these papers), and for his persistence in contacting the Editor in Chief of Youth and Society. I am also grateful for tools like Statcheck, which enabled me to very quickly spot some of the problems with these papers.

Friday, November 1, 2019

To summarize, for the moment, my series on the Zhang lab's strange media violence research

It never hurt to keep something of a cumulative record of one's activities when investigating any phenomenon, including secondary analyses.

In the case of the work produced in the lab of Qian Zhang, I have been trying to understand their work, and what appears to have gone wrong with their reporting, for some time. Unbeknownst to me at the time in 2014, I was already encountering one of the lab's papers when by the luck of the draw I was asked to review a manuscript that I would later find was coauthored by Zhang. As I have previously noticed, that paper had a lot of problems and I recommended as constructively as I could that the paper not be published. It was published anyway.

More explicitly, I found a weapons priming article published in Personality and Individual Differences at the start of 2016. It was an empirical study and one that fit the inclusion criteria for a meta-analysis that I was working on at the time. However, I ran into some really odd statistical reporting, leaving me unsure as to what I should use to estimate an effect size. So I sent what I thought was a very polite email to the corresponding author and heard nothing. After a lot of head-scratching, I figured out a way to extract effect size estimates that I felt semi-comfortable with. In essence the authors had no main effect for weapon primes on aggressive thoughts - and it showed in the effect size estimate and confidence intervals. That study really had a minimal impact on the overall mean effect size for weapon primes on aggressive cognitive outcomes in my meta-analysis. I ran analyses and later re-ran analyses and went on with my life.

I probably saw a tweet by Joe Hilgard who was reporting some oddities in another Zhang et al paper sometime in the spring of 2018. That got me wondering what all I was missing. I made a few notes, bookmarked what I needed to bookmark, and came back to that question a bit later in the summer of 2018 when I had a bit of time and breathing room. By this point I could comb through the usual archives, EBSCO databases, ResearchGate, and Google Scholar, and was able to hone in on a fairly small set of English-language empirical articles coauthored by Qian Zhang of Southwest University. I saved all the PDF files, and did something that I am unsure if anyone had done already: I ran the articles through statcheck. With one exception at the time, all the papers I ran through statcheck that had the necessary elements reported (test stat value, p-value, degrees of freedom) showed serious decision errors. In other words, the conclusions the authors were drawing in these articles were patently false based on what they had reported. I was also able to document that the reported degrees of freedom were inconsistent within articles, and often much smaller than the reported sample sizes. There were some very strange tables in many of these articles that presumably reported means and standard deviations but looked more like poorly constructed ANOVA summary tables.

I first began tweeting about what I was finding in mid-to-late September 2018. I think between some conversations via Twitter and email, I at least was convinced that I had spotted something odd, and that my conclusions so far as they went were accurate. Joe Hilgard was especially helpful in confirming what I had found, and then going well beyond that. Someone else honed in on inaccuracies in the reporting of the number of reaction time trials reported in this body of articles. So that went on throughout the fall of 2018. By this juncture, there were a few folks tweeting and retweeting about this lab's troubling body of work, some of these issues were documented by individuals in PubPeer, and editors were being contacted, with varying degrees of success.

By spring of this year, the first corrections were published - one in Youth and Society and a corrigendum in Personality and Individual Differences. To what extent those corrections can be trusted is still an open question. At that point, I began blogging my findings and concerns here, in addition to the occasion tweet.

This summer, a new batch of errata were made public concerning articles published in journals hosted by a publisher called Scientific Research. Needless to say, once I became aware of these errata, I downloaded those and examined them. That has consumed a lot of space on this blog since. As you are now well aware, these errata themselves require errata.

I think I have been clear about my motivation throughout. Something looked wrong. I used some tools now at my disposal to test my hunch and found that my hunch appeared to be correct. I then communicated with others who are stakeholders in aggression research, as we depend on the accuracy of the work of our fellow researchers in order to get to as close an approximation of the truth as is humanly possible. At the end of the day, that is the bottom line - to be able to trust that the results in front of me are a close approximation of the truth. If they are not, then something has to be done. If authors won't cooperate, maybe editors will. If editors don't cooperate, then there is always a bit of public agitation to try to shake things up. In a sense, maybe my role in this unfolding series of events is to have started a conversation by documenting what I could about some articles that appeared to be problematic. If the published record is made more accurate - however that must occur - I will be satisfied with the small part I was able to play in the process. Data sleuthing, and the follow-up work required in the process, is time-consuming and really cannot be done alone.

One other thing to note - I have only searched for English-language articles published by Qian Zhang's lab. I do not read or speak Mandarin, so I may well be missing out on a number of potentially problematic articles in Chinese-language psychological journals. If someone who does know of such articles wishes to contact me please do. I leave my DM open on Twitter for a reason. I would especially be curious to know if there are any duplicate publications of data that we are not detecting. 

As noted before, how all this landed on my radar was really just the luck of the draw. A simple peer review roughly five years ago, and a weird weapons priming article that I read almost four years ago were what set these events in motion. Maybe I would have noticed something was off regardless. After all, this lab's work is in my particular wheelhouse. Maybe I would not have. Hard to say. All water under the bridge now. What is left is what I suspect will be a collective effort to get these articles properly corrected or retracted.

Monday, October 28, 2019

Consistency counts for something, right? Zhang et al. (2019)

If you manage to stumble upon this Zhang et al. (2019) paper, published in Aggressive Behavior, you'll notice that this lab really loves to use a variation of the Stroop Task. Nothing wrong with that in and of itself. It is, after all, presumably one of several ways to attempt to measure the accessibility of aggressive cognition. One can get mean differences between reactions times (rt) for aggressive words and for nonaggressive words under different priming conditions and see if the stimuli with what we believe is violent content make aggressive thoughts more accessible - in this case with reactions times being higher for aggressive words than nonaggressive words (hence, higher positive difference scores). I don't really want to get you too much into the weeds, but I just think having that context is useful in this instance.

So far so good, yeah?

Not so fast. Usually the differences we find in rt between aggressive and nonaggressive words in these various tasks - including the Stroop Task - are very small. We're talking maybe single digit or small double digit differences in milliseconds. As has been the case with several other studies where Zhang and colleagues have had to publish errata, that's not quite what happens here. Joe Hilgard certainly noticed (see his note in PubPeer). Take a peek for yourself:


Hilgard notes another oddity as well as the general tendency for the primary author (Qian Zhang) to essentially stonewall requests for data. This is yet another paper I would be hesitant to cite without access to data, given that this lab already has an interesting publishing history, including some very error-prone errata for several papers published from this decade.

Note that I am only commenting very briefly on the cognitive outcomes. The authors also have data analyzed using a competitive reaction time task. Maybe I'll comment more about that at a later date.

As always, reader beware.

Reference:

Zhang, Q., Cao, Y., Gao, J., Yang, X., Rost, D. H., Cheng, G., Teng, Z., & Espelage, D. L. (2019). Effects of cartoon violence on aggressive thoughts and aggressive behaviors. Aggressive Behavior, 45, 489-497. doi: 10.1002/ab.21836

Sunday, October 27, 2019

Postscript to the preceding

I am under the impression that the body of errata and corrigenda from the Zhang lab were composed as hastily and without care as were the original articles themselves. I wonder how much scrutiny the editorial teams of these respective journals gave these corrections as they were submitted. I worry that little scrutiny was involved, and it is a shame that once more post-peer-review scrutiny is all that is available.

Erratum to Zhang, Zhang, & Wang (2013) has errors

This is a follow up to my commentary on the following paper:

Zhang, Q. , Zhang, D. & Wang, L. (2013). Is Aggressive Trait Responsible for Violence? Priming Effects of Aggressive Words and Violent Movies. Psychology, 4, 96-100. doi: 10.4236/psych.2013.42013

The erratum can be found here.

It is disheartening when an erratum ends up being more problematic than the original published article. One thing that struck me immediately is that the authors continue to insist that they ran a MANCOVA. As I stated previously:
It is unclear just how a MANCOVA would be appropriate as the only DV that the authors consider for the remaining analyses is a difference score. MANOVA and MANCOVA are appropriate analytic techniques for situations in which multiple DVs are analyzed simultaneously. The authors fail to list a covariate. Maybe it is gender? Hard to say. Without an adequate explanation, we as readers are left to guess. Even if a MANCOVA were appropriate, Table 4 is a case study in how not to set up a MANCOVA table. Authors should be explicit about what they are doing as possible. I can read Method and Results sections just fine, thank you. I cannot, however, read minds.

In essence, my initial complaint remains unaddressed.  One change, Table 4 is now Table 1, and it has different numbers in it. Great. I still have no idea (nor would any reasonably-minded reader), based on the description given, what the authors used as a covariate nor do I know what purported multiple DVs were used simultaneously. This is not an analysis I use very often in my own work, although I have certainly done so in the past. I do have an idea of how MANOVA and MANCOVA tables would be set up, and how those analyses would be described. I did a fair amount of that for my first year project at Mizzou a long time ago. The authors used as their DV a difference score (diff between RT aggressive words vs RT nonaggressive words), which would rule out the need for a MANOVA. And since no covariate is specified, a MANCOVA would be ruled out. I am going to make a wild guess that the partial summary table that comprises Table 1 will end up being nonsensical as have been similar tables generated in papers by this lab, including errata and corrigenda. I don't expect to be able to generate the necessary error MS, which I could then use to estimate the pooled SD.

I also want to note that the description of Table 2 as characterized by the authors and the numbers in Table 2 do not match up. I find that troubling. I am assuming that the authors mislabeled the columns, and intended for the low trait and high trait columns to be reversed. It is still sloppy.
At least when I ran this document through Statcheck, the findings, as reported, appeared clean - no inconsistencies and no decision inconsistencies. I wish that provided cold comfort. Since I don't know if I can trust any of what I have read in either the original document or the current erratum, I am not sure that I there is any comfort to be had.

What saddens me is that so much media violence research is based on WEIRD samples. That influences the generalizability of the findings. That also limits the scope of any skepticism I and my peers might have about media violence effects. We need good non-WEIRD research. So the fact that there is a lab that is generating a lot of research that is non-WEIRD, but is riddled with errors is a major disappointment.

At this juncture, the only cold comfort I would find is if the lot of the problematic studies from this lab were retracted. I do not say that lightly. I view retraction as a last resort, when there is no reasonable way for the record to be corrected without removing the paper itself. Doing so appears to be necessary for at least a few reasons. One, meta-analysts might try to use this research - either the original article or the erratum (or both if they are not paying attention) to generate effect size estimates. If we cannot trust the effect size estimates we generate, it's pretty much game over. Two, given that in a globalized market we all consume much of the same media (or at least the same genres), it makes sense to have evidence from not only WEIRD samples but also non-WEIRD samples. Some of us might try to understand just how violent media affect samples from non-WEIRD populations in order to understand if our understanding of these phenomena are universal. The findings generated from this paper and from this lab more broadly do not contribute to that understanding. If anything, the findings detract from our ability to get any closer to the truth. Three, the general public latches on to whatever seems real. If the findings are bogus - either due to gross incompetence or fraud - then the public is essentially being fleeced, which to me is simply unacceptable. The Chinese taxpayers deserved better. So do all of us who are global citizens.