Showing posts sorted by relevance for query zhang. Sort by date Show all posts
Showing posts sorted by relevance for query zhang. Sort by date Show all posts

Friday, December 6, 2019

Now about those Youth and Society retractions involving Qian Zhang

Hopefully you have had a moment to digest the recent article about the retraction of two of Qian Zhang's papers in Retraction Watch. I began tweeting about one of the articles in question around late September, 2018. You can follow this link to my tweet storm for the now retracted Zhang, Espelage, and Zhang (2018) paper. Under a pseudonym, I initially documented some concerns about both papers in PubPeer: Zhang, Espelage, and Zhang (2018) and Zhang, Espelage, and Rost (2018).

Really, what I did was to upload the papers into the online version of Statcheck, and flag any decision inconsistencies I noticed. I also tried to be mindful of any other oddities that seemed to stick out at that time. I might make note of df that seemed unusual given the reported sample size, for example, or problems with tables presuming to report means and standard deviations. By the time I would have looked at these two papers, I suspect that I was already concerned that papers from Zhang's lab showed a pervasive pattern of errors. Sadly, these two were no different.

With regard to Zhang, Espelage, and Zhang (2018), a Statcheck scan showed three decision errors. In this case, these were errors where the authors reported findings as statistically significant when they were not - given the test statistic value, the degrees of freedom, and the level of significance the authors tried to report.

The first decision inconsistency has to do with an assertion that playing violent video games increased accessibility of aggressive thoughts. The authors initially reported the effect as F(1, 51) = 2.87, p < .05. The actual p-value would have been .09634, according to Statcheck. In other words, there is no main effect for violent content of video games in this sample. Nor was a video game type by gender interaction found: F(1, 65) = 3.58, p < .01 actual p-value: p = 0.06293. Finally, there is no game type by age interaction: F(1, 64) = 3.64, p < .05 actual p-value: p = 0.06090. Stranger still, the sample was approximately 3000 students. Why were the denominator degrees of freedom so small for these reported test statistics? Something did not add up. Table 1 from Zhang, Espelage, and Zhang (2018) was also completely impossible to interpret - an issue I have highlighted in other papers that have been published from his lab:

A few months later, a correction would be published in which the authors would purport to correct a number of errors found on several of the pages of the original article, as well as Table 1. That was wonderful insofar as it went. However, there was a new oddity. The authors purported to only use 500 of the 3000 participants in order to have a "true experiment" - which was one of the more interesting uses of that term I have read over the course of my career. And as Joe Hilgard has aptly noticed, problems with the descriptive statistics continued to be pervasive - implausible and impossible cell means and marginal means and standard deviations, for example.

With regard to the Zhang, Espelage, and Rost (2018) article, my initial flag was simply for some decision errors in Study 1, in which the authors were attempting to establish that their stimulus materials were equivalent across a number of variables except for the level of violent content, and consistent across subsamples, such as sex of participant (male/female). Given the difficulty that exists in obtaining, say, film clips that are sufficiently equivalent except for level of violence, due diligence in using materials that are as equivalent as possible, except for the IV or IVs is to be admired. Unfortunately, There were several decision errors that I flagged after a Statcheck run.

As noted at the time, contra the authors' assertion, there was evidence that there were some rated differences between the violent film (Street Fighter) and the nonviolent film (Air Crisis) in terms of pleasantness - t(798) = 2.32, p > .05 actual p-value: p = 0.02059 - and fear - t(798) = 2.13, p > .05 actual p-value: p = 0.03348  To the extent that failure to control for those factors might have impacted subsequent analyses in Study 2 is of course debatable. It is clear that the authors cannot demonstrate, based on their reported analyses, that they had films that were equivalent on variables that they identified as important to hold constant with only violent content varying. The final decision inconsistency suggested that there was a sex difference in ratings of the variable fear, contrary to authors claim, t(798) = -2.14, p > .05 actual p-value: p = 0.03266. How much that impacted the experiment in study 2 was not something I thought I could assess, but I found it troubling and worth flagging. At minimum, the film clips were less equivalent than reported, and the subsamples were potentially reacting differently to these film clips than reported.

Although I did not comment on Study 2, Hilgard demonstrated that there was a consistency in the pattern of reported means in this paper were strikingly similar to the pattern of means reported in a couple of earlier papers in which Zhang was a lead or coauthor in 2013. That is troubling. If you have followed some of my coverage of Zhang's work on my blog, you are well aware that I have actually discovered at least one instance in which a reported test statistic was directly copied and pasted from one paper to another. Make of it what you will. Dr. Hilgard was able to eventually get a hold of the data that were to accompany a correction to that article, and as noted in the coverage in Retraction Watch, the data and analyses were fatally flawed.

I was noticing a pervasive pattern of errors in these papers, along with others I was reading by Zhang and colleagues at the time. These are the first papers on which Zhang is a lead or coauthor to be retracted. I am willing to bet that these will not be the last, given the evidence I have been sharing with you all here and on Twitter over the last year. I have already probably stated this repeatedly about these retractions - I am relieved. There is no joy to be had here. This has been a bad week for the authors involved. Also please note that I am taking great care here not to assign motive. I think that the evidence speaks for itself that the research was poorly conducted and poorly analyzed. That can happen for any of a number of reasons. I don't know any of the authors involved. I have some awareness of Dr. Espelage's work in bullying, but that is a bit outside my own specialty area. My impression of her work has always been favorable, and these retractions notwithstanding, I see no reason to change my impression of her work on bullying.

If I was sounding alarms in 2018 and onward, it is because Zhang had begun to increasingly enlist as collaborators well-regarded American and European researchers, and was beginning to publish in top-tier journals in various specialties within the Psychological Sciences, such as child and adolescent development and aggression. Given that I thought a reasonable case could be made that Zhang's reputation for well-conducted and analyzed research was far from ideal, I did not want to see otherwise reputable researchers put their careers on the line. My fears to a certain degree are now being realized.

Note that in the preparation of this post, I relied heavily on my tweets from Sept. 24, 2018 and a couple posts I published pseudonymously in PubPeer (see links above). And credit where it is due. I am glad I could get a conversation started about these papers (and others) by this lab. Joe Hilgard has clearly put a great deal of effort and talent into clearing the record since. Really we owe him a debt of gratitude. And also a debt of gratitude to those who have asked questions on Twitter, retweeted, and refused to let up on the pressure. Science is not self-correcting. It takes people who care to actively do the correcting.

Friday, July 29, 2022

A Blast From the Past: Retractions and Meta-Analysis Edition

I stumbled across this article, Media and aggression research retracted under scrutiny, and found it to be an interesting short read. The article's author chronicles some recent retractions, and what had been another on-going investigation of several papers coauthored by Qian Zhang of Southwest University. I've written enough about his work over the last few years. I think referring to many of Zhang's papers having "been called into question" is a fair assessment. 

Part of the story chronicles Samuel West, who included one of Zhang's papers in his meta-analysis at the request of a reviewer. His meta-analysis would undergo another round of peer review around the time he learned of that particular Zhang paper being under investigation at the same journal. Ouch. West certainly has legitimate concerns about including a potentially dodgy finding in his meta-analysis. In this case, the paper by Zhang and colleagues was not retracted, but I am sure West has his misgivings about including the paper in his database in the first place. I can certainly empathize. My most recent published meta-analysis included one of Zhang's papers that would eventually get retracted early this year. That said, there are plenty of papers generated from Zhang's lab with obvious problems, or, in the case of his more recent work, have problems that are more cleverly hidden. I agree with Amy Orben that the fact that problematic studies continue to remain in journals and meta-analyses is "a major problem" when we think about how politicized media violence research is. Requiring archiving of data, data analyses, and research protocols probably helps to the extent that it is required - at least anything that might be incorrect or fraudulent can more easily be sniffed out. Otherwise, one can only hope for sleuths with enough time on their hands and no concerns for career repercussions for blowing the whistle on published papers that should have never seen the light of day. Good luck with that.

I do take issue with Zhang's characterization of Hilgard as someone who is "just trying to make his name based just on claiming that everyone else does bad research." I get that Zhang is a bit sore about the retractions, and Hilgard was the person who contacted Zhang and a plethora of journal editors regarding the papers in question. That said, there was plenty of chatter about Zhang's work in 2018 and onward, and there were probably several of us who just wanted to know that we hadn't gone insane, and that the obvious data errors, including degrees of freedom that were inaccurate, means and standard deviations that were mathematically impossible, and tables that made no sense really were what we thought they were. Hilgard was far and away better connected to the sphere of media violence research as an active researcher himself, and had the data analytic know-how and the connections that come with being at a R-1 university to do what needed to be done. Aside from that, Hilgard made plenty of positive contributions to the methodology side of psychological science, and from interacting with him online and in person over the years, I'll simply say he's a good person to know. 

I think this article is somewhat helpful in pointing out that even those who believe there is a link between violent content in media (such as video games) and aggression can view Zhang's work and see it for what it is, and express an appropriate level of skepticism. At the end of the day, one can take a philosophical perspective that there is "no one right way to look at the data" and that's all well and good. But at the end of the day, if the analyses show decision errors, and the means and standard deviations forming the basis for those analyses are simply mathematically impossible, the only reasonable conclusion that can be made is that the data and analyses in their present form cannot be accepted as valid. 

The only bone I really have to pick is that the author characterizes the body of media violence research as asking the question of whether or not "violent entertainment causes violence". Although I am aware that there are researchers in this area of inquiry who would draw that conclusion, there are plenty of other investigators who view what we can learn based on our available methods much more cautiously (a lot of aggression is mild, after all). There are also plenty of skeptics who doubt that there is any link between media violence and even the mild forms of aggression that we can measure. As far as I am aware, there is no link between exposure to violent content in mass media and violent behavior in everyday life. All that said, this is a useful article that captures a series of events that I know quite intimately. 

Suddenly, I am in the mood for some cartoon violence. I think I'll watch some early episodes of Rick and Morty. Goodnight.

Saturday, September 14, 2019

Prelude to the latest errata

Now that there have been some relatively new developments regarding research from Qian Zhang's lab, I think the best thing to do is to give you all some context before I proceed. So let's look at some of the blog posts I have composed about the articles that are now being presumably corrected:

A. Let's start out with the most recent and work our way backwards. First, let's travel back to the year 2016. You can easily find this paper, which is noteworthy for being submitted roughly a year or so after its fourth author had passed away.

Tian, J. , Zhang, Q. , Cao, J. and Rodkin, P. (2016). The Short-Term Effect of Online Violent Stimuli on Aggression. Open Journal of Medical Psychology, 5, 35-42. doi: 10.4236/ojmp.2016.52005

See these blog posts:

"And bad mistakes/I've made a few"*: another media violence experiment gone wrong

 Maybe replication is not always a good thing

It Doesn't Add Up: Postscript on Tian, Zhang, Cao, & Rodkin (2016)

A tale of two Stroop tasks

B. Now let's revisit the year 2014. There is one article of note here. I had one post on this article at the time, and had wished I had devoted a bit more time to it. Note that in many of these earlier articles, Zhang goes by Zhang Qian, and for whatever reason, the journal of record recommends citing Qian as the family name. Make of that what you will. Following is the reference.

Tian, J. & Zhang, Q. (2014). Are Boys More Aggressive than Girls after Playing Violent Computer Games Online? An Insight into an Emotional Stroop Task. Psychology, 5, 27-31. doi: 10.4236/psych.2014.51006.

See this blog post:

Funny, but sad

The year 2013 brings us two papers to consider. I only devoted a single blog post to the first article referenced. The second article got referenced twice as I noticed the same oddity when it came to the way the authors were describing the Stroop task and analyzing data based on that task.

C. First we will start here with a basic film violence study.

Zhang, Q. , Zhang, D. & Wang, L. (2013). Is Aggressive Trait Responsible for Violence? Priming Effects of Aggressive Words and Violent Movies. Psychology, 4, 96-100. doi: 10.4236/psych.2013.42013

See this blog post:

About those Stroop task findings (and other assorted oddities)

D. And here is the article in which the authors use the Stroop task in a most remarkably odd way.

Zhang, Q. , Xiong, D. and Tian, J. (2013) Impact of media violence on aggressive attitude for adolescents. Health, 5, 2156-2161. doi: 10.4236/health.2013.512294

See this blog post:

Some more oddness (Zhang, Xiong, & Tian, 2013)

I could probably add some other work for context, as there are some pervasive patterns that show up across studies over the course of this decade. As the authors have begun to rely upon larger data sets, there are some other troubling practices, such as using only a small fraction of a sample to analyze data (something I vehemently oppose as a practice). Whether the articles are published in low-impact journals or high-impact journals is of no importance in one sense: poorly conducted research is poorly conducted research, and if it needs to be corrected, it is up to the authors to do so in as transparent and forthright a manner as possible. That said, as this lab is getting work published in higher impact journals, the potential for incorrectly analyzed data and hence misleading findings to poison the proverbial well increases. That should trouble us all.

I want to end with something I said a few months ago, as it is important to understand where I am coming from as I once more proceed:
Although I don't have evidence that the Zhang lab was involved in any academic misconduct, and I have no intention of making accusations to that effect, I do think that some of the data reporting itself is at best indicative of incompetent reporting. All I can do is speculate, as I am unaware of anyone who has managed to actually look at this lab's data. What I can note is that there is a published record, and that there are a number of errors that appear across those published articles. Given the number of questions I think any reasonable reader of media violence research might have, Zhang and various members of his lab owe it to us to answer those questions and to provide us with the necessary data and protocols to accurately judge what went sideways.
The reason I emphasize this point is because this is really not personal. This is a matter of making sure that those of us at minimum who do study media violence research have accurate evidence at our disposal.

Friday, January 29, 2021

Research Confidential: How Self-Correcting is Science?

The question is arguably rhetorical. Science in and of itself is not self-correcting. It takes living, breathing human beings to notice something is wrong, take the time and make the effort to report what is wrong to relevant stakeholders (e.g., journal editors, relevant university adminstrations, etc.), and then have good reason to believe that the relevant stakeholders will show due diligence, correct or retract flawed papers as needed, and otherwise hold those responsible for the flaws, whether due to sheer incompetence or fraud, accountable. In an ideal world, that is how it would work. In this world, it's considerably more complicated, and often more than a bit disheartening.

If you are a regular reader of this blog, you are quite aware of our favorite media violence researcher who is notorious for some of the worst papers published in that particular niche area of psychology - Qian Zhang of Southwest University. I have documented, over the last couple years or so, some of the most insane tables with means, standard deviations, and test statistics that simply are impossible to interpret. I have reported test statistics that, based on the degrees of freedom reported, would have to be incorrect. I have reported discrepancies between degrees of freedom for test statistics and the sample size reported. I have documented evidence of potential plagiarism and self-plagiarism - the latter due to the tendency for Zhang to rely heavily on copying and pasting from one paper to another. I have also found some amusing typos that resulted from Zhang's tendency to copy and paste tables from paper to paper. I've tagged Zhang's work as I have documented here (for your convenience) and on PubPeer under a pseudonym. 

Dr. Joe Hilgard has gone considerable further than have I. He's blogged about his own experiences in documenting problems with Zhang's work in great detail, and the efforts he's made to contact journal editors along with officials at Zhang's university, offering painstaking evidence of the problems he has discovered. You can read about Joe Hilgard's efforts, and the decidedly mixed and disappointing outcome of his efforts here. You should really take to the time to read Hilgard's post as it is thorough and damning. The short version? Some journal editors responded rather well, and in one case very quickly to retract two papers that were clearly unsound. Other journal editors have either stonewalled or ignored Hilgard's concerns. Zhang's university cleared him of wrongdoing, chalking it all up to Zhang being "deficient in statistical knowledge and research methods." So in other words, the university writes it off as "the guy's merely an idiot, but hey, let's just give him a remedial stats course and call it even." I agree with Hilgard that the university's failure to take action is not that surprising, as universities seem to be in the business of taking care of their own, especially if the researcher in question might be bringing in grants or other forms of prestige. So the guy maybe fudges some numbers and has no idea what random assignment means. There's nothing to see here. Move along.

My take on the matter is that the most charitable view that one could take based on the body of Qian Zhang's work is that this is a researcher who is grossly incompetent, but that a more probably defensible case can be made that his activities are on some level fraudulent. I am more inclined to the latter less charitable view. I've seen too much. Regardless, this is research that should have never made it past peer review. I agree with Hilgard that this body of research is very problematic given that as long as it remains published, it will distort our understanding of what is actually happening with stimuli such as video games that contain violent content on outcome variables such as aggressive behavior or cognition. Meta-analyses are especially vulnerable given that some of the reported findings by Zhang rely on large samples. Those results could artificially inflate effect sizes, leading meta-analysts and those consuming meta-analyses to believe that an overall effect is stronger than it actually is. 

This is one of the dark alleys I mentioned a few years ago. And given what Hilgard has experienced and what I've experienced in my own way, it's one that few leave with any sense of hope for the state of this particular are of psychological inquiry. If blatantly problematic papers, ones where the problems are so obvious that a beginning methods student could discover them, cannot be retracted within a short window of time, what is going on with work in which potentially fraudulent data analyses are more cleverly presented? What else is out there that cannot be trusted? That is something that should cause us all to lose some sleep.

One final thought for anyone thinking of collaborating with Zhang: don't. If you absolutely cannot help yourself, insist on seeing the data before agreeing to be part of that particular project. I'd say that is a safe practice regardless of the situation. If I take on a statistician for a project, or someone who is at least better versed in a particular statistical method than I am, I insist on sending the data set or database, and I expect that the statistician on the project will double check my work and ask difficult questions as needed. That can save a lot of grief, assuming that the statistician involved is actually looking at what is being sent. One of the tragedies for some of Zhang's coauthors is that they've never had access to the data sets to which they lent their names and reputations, nor were they apparently allowed access. That is not how we do science, folks.

In the meantime, more papers are in the pipeline to be published by this particular author, and it will become more of a struggle to keep up with the dross that is likely to be found in any of those papers. Again, that is something that should cause us all to lose some sleep.

Thursday, October 8, 2020

Postscript to the preceding: This isn't the first odd mistake for Zhang

 Previously, I noted that the latest Zhang et al. (2020) paper had at least one serious error: that instead of computing a difference between reaction time for aggressive (weapon) images and neutral images, the authors used simply the reaction times to the weapon images as the DV. Hence, we as the readers are left with a misleading set of analyses and a potentially misleading narrative. Fortunately, the authors had already shared their data, which made detecting the error fairly easy. Why the reaction times for the neutral images and then the difference scores (which would have been the real DV) didn't have their own column is only something that the authors can answer.

Oftentimes, with this lab (as is probably the case with others), it is often difficult to glean whether or not variables are entered and computed correctly based on the information appearing in a published paper. Whether those omissions are a bit of sleight of hand or simple human error or misunderstanding is often difficult to deduce. However, sometimes authors make it easy for the readers to see for themselves that the authors have goofed. I have found some rather odd analyses in which IVs were not quite analyzed correctly as well as DVs.

One of my favorite papers published by the Zhang lab, just for the sheer madness it contained, was the on published in Personality and Individual Differences nearly five years ago. That was the first, and I think only, effort these authors made to replicate and extend research on the weapons priming effect (itself a fairly controversial topic). The DV situation appears okay in the initial analysis under section 4.1. However, where things fall apart (aside from a grossly undersized df, given sample size) was that the authors only examined the difference in reaction times between aggressive and neutral words under the weapon prime condition, while completely ignoring the neutral prime condition. The authors eventually did correct the df for that section in a pretty massive corrigendum. However, they never did address that they had done the wrong analysis in order to establish a weapons priming effect. They really should have read more carefully Anderson et al. (1998) in order to do so. The authors needed to establish that the difference between rts in the treatment and control conditions were larger, and in the predicted direction, when participants saw weapons than when they were presented with neutral images. Also left unanswered was the nagging question of the three-way interaction effect that was a duplicate of another three-way interaction effect in another paper authored by this same research team. I got the impression that the current editor in chief at Personality and Individual Differences was not much in the mood for dealing with this mess to begin with, and that any superficial corrections were extracted from Zhang et al. (2016) was probably a minor miracle. In theory, since the authors changed a single digit in the F-test for the three-way interaction, perhaps the point is now moot. I am still concerned that a certain amount of self-plagiarism happened, but the editor-in-chief chose to let it go. As was the case with the most recent article in question, the Zhang lab had enlisted an established American aggression researcher, Phillip Rodkin. Rodkin's wheelhouse was more in the area of bullying, and not so much media violence, so this seemed like an odd choice for a collaborator for a media violence paper. I honestly don't know how much access Rodkin had to the original data, nor could I comment on whether he would have known what to look for when checking out the analyses. He had already been deceased for a while when this paper was published. Hence, we will likely never know.

The Zhang et al. (2016) paper shared something strikingly in common with a paper in which Zhang was second author, and Rodkin also was a collaborator. There was a three-way interaction that was deemed nonsignificant in each paper, although according to a Statcheck analysis, the three-way interaction would have to have been statistically significant based on what was originally reported in each paper. Publication of duplicate analyses is presumably serious business, but apparently the powers that be can overlook such matters. Perhaps the corrigendum on the Zhang et al (2016) paper makes the point moot, as I noted earlier. The erratum in the other paper entirely ignores the pesky issue of that three-way interaction effect. 

As I have probably said too many times, I find this state of affairs to be very disappointing. As someone who still finds media violence research interesting (although definitely from the standpoint of a skeptic), I treasure efforts by researchers who study non-WEIRD populations. As an educator and researcher who is very eager to decolonize my particular areas of expertise, I would ordinarily welcome work coming out of China. Unfortunately, the work from this lab is so chock full of errors that it is best left uncited. Hold out for the real thing. Hold out for competently and ethically conducted work. 

References:

https://doi.org/10.1016/j.paid.2015.09.017 

https://doi.org/10.1016/j.paid.2018.12.010 

http://dx.doi.org/10.4236/ojmp.2016.52005

http://dx.doi.org/10.4236/ojmp.2019.83005

Sunday, April 7, 2019

Closing the books on a correction (Zhang et al., 2016)

When I was updating the weapons effect database for a then-in-progress meta-analysis a little over three years ago, I ran across a paper by Zhang, Tian, Cao, Zhang, & Rodkin (2016). You can read the original here, as it required significant corrections. The Corrigendum can be found here.

Initially, I was excited, as it is not often that one finds a weapons effect paper published that is based on non-American or non-European samples. There were obvious problems from the start. First, although the authors purport to measure aggression in adolescents (in reality the sample were pre-adolescent children), in reality the dependent variable was a difference in reaction time between aggressive and non-aggressive words. To put it another way, the authors were merely measuring accessibility of aggressive thoughts that presumably would be primed by mere exposure to weapons.

The analyses themselves never quite added up, which made determining an accurate effect size estimate from their work to be, shall we say, a wee bit challenging. I attempted to contact the corresponding author asking for data and any code or syntax used in the hopes of reproducing the analyses and getting the information necessary and obtaining the effect size estimate that would most closely approximate the truth. That email was sent on January 26, 2016. I never heard from Qian Zhang. I figured out a work-around in order to obtain a satisfactory-enough effect size estimate and moved in.

But that paper always bothered me once the initial excitement wore off. I am well aware that I am far from alone in having some serious questions about the Zhang et al. (2016) article. Some of those could be written off as potential typos: there were some weird discrepancies in degrees of freedom across the analyses. The authors contended that they established that they had replicated work I had been involved in conducting (Anderson, Benjamin, & Bartholow, 1998) by simply examining if reaction times to aggressive words were more rapid when primed with weapons than neutral images. In our experiments, we used the difference between aggressive and non-aggressive words as our dependent variable. And based on the degrees of freedom reported, it appeared that the analysis was based on one subsample, as opposed to the complete sample. So obviously there are some red flags.

The various subsample analyses using a proper difference score (they call it AAS) also looked a bit off. And of course the MANOVA table seemed unusual, especially since the unit of analysis appeared to be their difference score (reaction times for aggressive words minus non-aggressive words) - a single dependent variable - as opposed to multiple dependent variables. Although I have rarely used MANOVA and am unlikely to use MANOVA in my own research, I certainly had enough training to know what such analyses should look like. My understanding is that one would report MS, df, and F values for each IV-DV relationship, with the understanding that there will be at least two DVs for every IV. A cursory glance at the most recent edition I had of a classic textbook on multivariate statistics by Tabachnick and Fidell (2012) convinced me that the summary table reported in the article was inappropriate, and would confuse readers rather than enlighten them. There were other questions about the extent to which the authors more or less copied and pasted content from the Buss and Perry (1992) article in which they present their Aggression Questionnaire. Those as of yet have not been adequately addressed, and I suspect they never will.

So, I ran the analyses the authors provided in statcheck.io. I had even more questions. There were numerous errors, including decision errors even assuming that the test statistics and their respective degrees of freedom were accurate. Just to give you a flavor, here are my initial statcheck analyses:



As you can see, the authors misreport F(1, 155) = 1.75 p < .05 (actual p = .188), F(1, 288) = 3.76 p < .01 (actual p = .054), and F(1, 244) = 1.67, p < .05 (actual p = .197). The authors also appeared to misreport a three-way interaction as non-significant that clearly was statistically significant. Statcheck could not catch that one due to the authors' failure to include any degrees of freedom in their report. Basically, there was no good reason to trust the analyses at this point. Keep in mind that what I have done here is something that anyone with a basic graduate-level grounding in data analysis and access to Statcheck could compute. Anyone can reproduce what I did. That said, communicating with others about my findings was comforting: I was not alone in seeing what was clearly wrong.

In consultation with some of my peers, something else jumped out: the authors reported an incorrect number of trials. The authors reported 36 primes and 50 goal words which were each randomly paired. The authors reported a total number of trials as 900. However, if you do the math, it becomes obvious that the actual number of trials was 1800. As someone who was once involved in conducting reaction time experiments, I know the importance of not only assessing the necessary number of trials depending on the number of stimuli and target words that must be randomly paired, but also the importance of accurately reporting the number of trials required of participants. It is possible that given their description in the article itself, the authors took the number 18 (for weapons, for example) and multiplied it by 50. In itself, that seems like a probable and honest error. It happens, although it would have been helpful for this sort of thing to have been worked out in the peer review process.

The corrections in the corrigendum suggest a rather massive correction to the article. The presumed MANOVA table never quite gets completely resolved to satisfaction, and a lingering decision error remains. The authors also start using the term marginally significant to refer to a subsample analysis that made me cringe. The concept of marginal significance was supposed to have been swept into the dustbin of history a long time ago. We are well enough along into the 21st century to avoid that vain attempt to rescue a finding altogether.  Whether the corrections noted in the corrigendum are sufficient to save the conclusions the authors wished to make in the article is questionable. At minimum, we can conclude that Zhang et al. (2016) did not find evidence of weapon pictures priming aggressive thoughts, and even their effort to base a partial replication on subsample analyses was not sufficient. It is a non-replication, plain and simple.

My recommendation is not to cite Zhang et al. (2016) unless absolutely necessary. If one is conducting a relevant meta-analysis, citation is probably unavoidable. Otherwise, the article is probably worth citing if one is writing about questionable reporting of research, or perhaps as an example of research that fails to replicate a weapons priming effect.

Please note that the intention is not to attack this set of researchers. My concern is strictly on the research report itself, and the apparent inaccuracies contained in the original research report. I am quite pleased that however it transpired, the editor and authors were able to quickly make corrections in this instance. Mistakes get made. The point is to make an effort to fix them when they are noticed. That should at least be normal science. So kudos to those involved in making the effort to do the right thing here.

References

Anderson, C. A., Benjamin, A. J., Jr., & Bartholow, B. D. (1998). Does the gun pull the trigger? Automatic priming effects of weapon pictures and weapon names. Psychological Science, 9, 308-314. doi: 10.1111/1467-9280.00061

Buss, A. H., & Perry, M. (1992). The aggression questionnaire. Journal of Personality and Social Psychology, 63, 452-459. doi:10.1037/0022-3514.63.3.452

Tabachnick, B. G., & Fidell, L. S. (2012). Using multivariate statistics. New York: Pearson.

Zhang, Q, Tian, J., Cao, J., Zhang, D., & Rodkin, P. (2016). Exposure to weapon pictures and subsequent aggression in adolescence. Personality and Individual Differences, 90, 113-118. doi: 10.1016/j.paid.2015.09.017.

Sunday, October 13, 2019

Back to that latest batch of errata: Tian et al (2016)

So a few weeks ago I noted that there had been four recent corrections to papers published out of the Zhang lab. It's time to turn to a paper with a fun history to it:

Tian, J. , Zhang, Q. , Cao, J. and Rodkin, P. (2016). The Short-Term Effect of Online Violent Stimuli on Aggression. Open Journal of Medical Psychology, 5, 35-42. doi: 10.4236/ojmp.2016.52005

What really caught initially was that there was a 3-way interaction reported as nonsignificant in this particular article that was identical to the analysis of a similar 3-way interaction in another article published by Zhang, Tian, Cao, Zhang, and Rodkin (2016). Same numbers, same failure to report degrees of freedom, and the same decision error in each paper. Quite the coincidence, as I noticed before. Eventually, Zhang et al. (2016) did manage to change the numbers on several analyses on the paper published in Personality and Individual Differences. See the corrigendum for yourself. Even the 3-way interaction got "corrected" so that it no longer appeared significant. We even get - gasp - degrees of freedom! Not so with Tian et al. (2016) in OJMP. I guess that is the hill these authors will choose to die on? Alrighty then.

So if you want to really see what gets changed from the original Tian et al. (2016) paper, read here. Compared to the original, it appears that the decision errors go away - except of course for that pesky three-way ANOVA, which I guess the authors simply chose not to address. Gone to is any reference to computing MANCOVAs, which is what I would minimally expect, given that there was no evidence that such analyses were ever done - no mention of a covariate, nor any mention of multiple dependent variables to be analyzed simultaneously. This is at least a bit better. The table of means at least on the surface seems to add up. The new Table 1 is a bit funky. I've noticed with another one of the papers that the Mean Square error based on the Mean Square information for the main effect and interaction effects that the authors were interested in would not give an estimate of Mean Square error that would support the SDs supplied in the descriptive stats. That appears to be the case with this correction as well, to the extent that one can make an educated guess about Mean Square error based on an incomplete summary table. Even with those disadvantages in papers by other authors in the past, I have generally managed to get a reasonably close estimate of MSE, and hence with some simple computations estimate the pooled standard deviation. That I am unable to do so satisfactorily here is troubling. When I ran into this difficulty with another one of the corrections from this lab, I consulted with a colleague who quickly made it clear that the likely correct pooled MSE would not support the descriptive statistics as reported. So at least I am reasonably certain here that I am not making a mistake.

I also find it odd that the authors now discuss that viewing violent stimuli had no change on aggressive personality - as if the Buss and Perry Aggressiveness Questionnaire, which measures a stable trait would ever be changed by short term exposure to a stimulus like a brief clip of a violent film. What the authors might have been trying to state is that there was no interaction of scores on the Buss and Perry AQ and movie violence. That is only a guess in my part.

These corrections, as they are billed, strike me as very rushed, and potentially as mistake-ridden as the original articles. This is the second correction out of this new batch that I have had time to review and it is as problematic as the first. Reader beware.

Monday, October 28, 2019

Consistency counts for something, right? Zhang et al. (2019)

If you manage to stumble upon this Zhang et al. (2019) paper, published in Aggressive Behavior, you'll notice that this lab really loves to use a variation of the Stroop Task. Nothing wrong with that in and of itself. It is, after all, presumably one of several ways to attempt to measure the accessibility of aggressive cognition. One can get mean differences between reactions times (rt) for aggressive words and for nonaggressive words under different priming conditions and see if the stimuli with what we believe is violent content make aggressive thoughts more accessible - in this case with reactions times being higher for aggressive words than nonaggressive words (hence, higher positive difference scores). I don't really want to get you too much into the weeds, but I just think having that context is useful in this instance.

So far so good, yeah?

Not so fast. Usually the differences we find in rt between aggressive and nonaggressive words in these various tasks - including the Stroop Task - are very small. We're talking maybe single digit or small double digit differences in milliseconds. As has been the case with several other studies where Zhang and colleagues have had to publish errata, that's not quite what happens here. Joe Hilgard certainly noticed (see his note in PubPeer). Take a peek for yourself:


Hilgard notes another oddity as well as the general tendency for the primary author (Qian Zhang) to essentially stonewall requests for data. This is yet another paper I would be hesitant to cite without access to data, given that this lab already has an interesting publishing history, including some very error-prone errata for several papers published from this decade.

Note that I am only commenting very briefly on the cognitive outcomes. The authors also have data analyzed using a competitive reaction time task. Maybe I'll comment more about that at a later date.

As always, reader beware.

Reference:

Zhang, Q., Cao, Y., Gao, J., Yang, X., Rost, D. H., Cheng, G., Teng, Z., & Espelage, D. L. (2019). Effects of cartoon violence on aggressive thoughts and aggressive behaviors. Aggressive Behavior, 45, 489-497. doi: 10.1002/ab.21836

Tuesday, April 30, 2019

A postscript of sorts - research ethics and the Zhang lab

Although I don't have evidence that the Zhang lab was involved in any academic misconduct, and I have no intention of making accusations to that effect, I do think that some of the data reporting itself is at best indicative of incompetent reporting. All I can do is speculate, as I am unaware of anyone who has managed to actually look at this lab's data. What I can note is that there is a published record, and that there are a number of errors that appear across those published articles. Given the number of questions I think any reasonable reader of media violence research might have, Zhang and various members of his lab owe it to us to answer those questions and to provide us with the necessary data and protocols to accurately judge what went sideways.

The Ministry of Education of the People's Republic of China has its set of guidelines for what might constitute academic misconduct and the process involved in opening and conducting an investigation. Accordingly, Southwest University has its own set of policies regarding research ethics and investigations of potential misconduct. The university has a committee charged with that responsibility. What I am hoping is that enough readers who follow my blog will reach out to those in charge and demand answers, demand accountability. I think a very good case can be made that there are some problem papers authored by members of this lab in need of some form of correction, in whatever form that might take. How that gets handled is ultimately up to what Southwest University finds and to the editors themselves who are in charge of journals that have published the research in question. That will help all of us collectively, assuming Zhang's university and the editors of the journals who published his work do the right thing here.

Science is not self-correcting. It is dependent on conscious human effort to detect mistakes, reach out to others, give them a chance to respond, and failing that continue to agitate for corrections. When errors are found, it is up to those responsible for those errors to step up and take responsibility - and to do what is necessary to make things right. We should not have to hope that someone with just the right expertise is willing to make some noise in an academic environment that is essentially hostile to whistle-blowers, and keep making noise until something finally happens. That is time consuming and ultimately draining.

Once I saw Zhang and colleagues' weapons priming paper, published in PAID in 2016, I could not unsee it. There were questions about that paper that continued to gnaw at me, and led me to read more and ask more questions. Before long, I was neck-deep in a veritable swamp of error-ridden articles published out of that lab. I take some cold comfort in knowing I am not alone, and that there are others with far better methodological skills and reach than I possess who can make some noise as well. Bottom line: something is broken and I want it fixed. Make some noise.

Friday, November 1, 2019

To summarize, for the moment, my series on the Zhang lab's strange media violence research

It never hurt to keep something of a cumulative record of one's activities when investigating any phenomenon, including secondary analyses.

In the case of the work produced in the lab of Qian Zhang, I have been trying to understand their work, and what appears to have gone wrong with their reporting, for some time. Unbeknownst to me at the time in 2014, I was already encountering one of the lab's papers when by the luck of the draw I was asked to review a manuscript that I would later find was coauthored by Zhang. As I have previously noticed, that paper had a lot of problems and I recommended as constructively as I could that the paper not be published. It was published anyway.

More explicitly, I found a weapons priming article published in Personality and Individual Differences at the start of 2016. It was an empirical study and one that fit the inclusion criteria for a meta-analysis that I was working on at the time. However, I ran into some really odd statistical reporting, leaving me unsure as to what I should use to estimate an effect size. So I sent what I thought was a very polite email to the corresponding author and heard nothing. After a lot of head-scratching, I figured out a way to extract effect size estimates that I felt semi-comfortable with. In essence the authors had no main effect for weapon primes on aggressive thoughts - and it showed in the effect size estimate and confidence intervals. That study really had a minimal impact on the overall mean effect size for weapon primes on aggressive cognitive outcomes in my meta-analysis. I ran analyses and later re-ran analyses and went on with my life.

I probably saw a tweet by Joe Hilgard who was reporting some oddities in another Zhang et al paper sometime in the spring of 2018. That got me wondering what all I was missing. I made a few notes, bookmarked what I needed to bookmark, and came back to that question a bit later in the summer of 2018 when I had a bit of time and breathing room. By this point I could comb through the usual archives, EBSCO databases, ResearchGate, and Google Scholar, and was able to hone in on a fairly small set of English-language empirical articles coauthored by Qian Zhang of Southwest University. I saved all the PDF files, and did something that I am unsure if anyone had done already: I ran the articles through statcheck. With one exception at the time, all the papers I ran through statcheck that had the necessary elements reported (test stat value, p-value, degrees of freedom) showed serious decision errors. In other words, the conclusions the authors were drawing in these articles were patently false based on what they had reported. I was also able to document that the reported degrees of freedom were inconsistent within articles, and often much smaller than the reported sample sizes. There were some very strange tables in many of these articles that presumably reported means and standard deviations but looked more like poorly constructed ANOVA summary tables.

I first began tweeting about what I was finding in mid-to-late September 2018. I think between some conversations via Twitter and email, I at least was convinced that I had spotted something odd, and that my conclusions so far as they went were accurate. Joe Hilgard was especially helpful in confirming what I had found, and then going well beyond that. Someone else honed in on inaccuracies in the reporting of the number of reaction time trials reported in this body of articles. So that went on throughout the fall of 2018. By this juncture, there were a few folks tweeting and retweeting about this lab's troubling body of work, some of these issues were documented by individuals in PubPeer, and editors were being contacted, with varying degrees of success.

By spring of this year, the first corrections were published - one in Youth and Society and a corrigendum in Personality and Individual Differences. To what extent those corrections can be trusted is still an open question. At that point, I began blogging my findings and concerns here, in addition to the occasion tweet.

This summer, a new batch of errata were made public concerning articles published in journals hosted by a publisher called Scientific Research. Needless to say, once I became aware of these errata, I downloaded those and examined them. That has consumed a lot of space on this blog since. As you are now well aware, these errata themselves require errata.

I think I have been clear about my motivation throughout. Something looked wrong. I used some tools now at my disposal to test my hunch and found that my hunch appeared to be correct. I then communicated with others who are stakeholders in aggression research, as we depend on the accuracy of the work of our fellow researchers in order to get to as close an approximation of the truth as is humanly possible. At the end of the day, that is the bottom line - to be able to trust that the results in front of me are a close approximation of the truth. If they are not, then something has to be done. If authors won't cooperate, maybe editors will. If editors don't cooperate, then there is always a bit of public agitation to try to shake things up. In a sense, maybe my role in this unfolding series of events is to have started a conversation by documenting what I could about some articles that appeared to be problematic. If the published record is made more accurate - however that must occur - I will be satisfied with the small part I was able to play in the process. Data sleuthing, and the follow-up work required in the process, is time-consuming and really cannot be done alone.

One other thing to note - I have only searched for English-language articles published by Qian Zhang's lab. I do not read or speak Mandarin, so I may well be missing out on a number of potentially problematic articles in Chinese-language psychological journals. If someone who does know of such articles wishes to contact me please do. I leave my DM open on Twitter for a reason. I would especially be curious to know if there are any duplicate publications of data that we are not detecting. 

As noted before, how all this landed on my radar was really just the luck of the draw. A simple peer review roughly five years ago, and a weird weapons priming article that I read almost four years ago were what set these events in motion. Maybe I would have noticed something was off regardless. After all, this lab's work is in my particular wheelhouse. Maybe I would not have. Hard to say. All water under the bridge now. What is left is what I suspect will be a collective effort to get these articles properly corrected or retracted.

Sunday, May 5, 2019

Some more oddness (Zhang, Xiong, Tian, 2013)

Here's another relevant oldie but goodie by Zhang's lab:



What is odd about this one - aside from Table 3, which is just a mess - is that there is no way that the simple effects F for Boys would be statistically significant given the degrees of freedom. Recall that there are 74 participants in the experiment. That chews up one df. Now let's look at the df for each comparison that would have to be computed, regardless of whether or not it was reported. There are three main effects and a DV which we'll simply note is the difference in reaction time between aggressive and non-aggressive words in a Stroop task. Main effect for movie type costs 1 df, main effect for gender costs another df, and main effect for trait aggressiveness as defined in the article costs 2 df. There are the interactions of movie type x gender (1 df), movie type x trait aggressiveness (2 df) and gender x trait aggressiveness (2 df). The three-way interaction will take another 2 df. That leaves us with 62 df. A significant F-test for simple main effects would need to be practically 4.00. The obtained F in Table 4 for Boys is well below that threshold. And yet the authors claim statistical significance for that particular subsample. The authors want to claim a significant movie type x gender interaction, but have no significant simple main effects. That does not compute, folks.

This, by the way, is one of the better articles from this lab in terms of - at least on the surface - data reporting.

One more thing to note: the authors improperly use terminology commonly used by aggression researchers incorrectly. They are not measuring aggressive attitude levels. They are measuring accessibility of aggressive cognition (or thoughts, in lay terms). I see that mistake enough times to acknowledge that it is hardly unique to Zhang's lab. It is just an annoyance, I suppose.

The article can be read here.

File under articles that probably should not be cited without adequate clarification from the authors.


Reference:

 Zhang, Q. , Xiong, D. and Tian, J. (2013) Impact of media violence on aggressive attitude for adolescents. Health, 5, 2156-2161. doi: 10.4236/health.2013.512294

Sunday, April 21, 2019

Maybe replication is not always a good thing

Check out these two photos. Notice the similarity?


The top image is from here. Since you'll probably be directed to the corrected version of the article, I will recommend going here to view the original, and also taking a moment to read through the Corrigendum. The Corrigendum is hardly ideal in this case, but seems to clear at least some of the wreckage. Moving on...

The second article comes from here.

So each article reports the findings suggesting a three-way interaction is not significant. In each case, the authors are wrong. I noted that with the weapon-priming article earlier.

But wait. There's more.

Notice that although each article is testing a different prime stimulus (the passage from the top image is one where weapons were primes, and the passage from the bottom is one where level of video game violence is the prime), samples representing different populations (youth in the passage for the weapons priming article and college students in the article examining violent video games as primes), and samples that differed at least somewhat in size, these authors miraculously obtain the same test statistic. A miracle, you say? Bull, I say. This is a case where it would be very helpful for the public to have access to data from each study, as there is reason to wonder how the same finding was obtained in each with the aforementioned differences duly noted.

There is something seriously rotten in the state of aggression research, dear readers. It is past time that we took notice, as there is a pervasive pattern of problems with published articles generated from this particular lab. If it were just one bum article, I could probably write it off as "mistakes were made" and let go. We're way beyond that. The real worry is that authors from this lab are getting in to collaborative relationships with other researchers in the US and EU who are generally reputable. I have to wonder how much their partners know of the problems that exist with their already existing published record. In some cases I wonder how their partners got chosen. The late Philip Rodkin was a researcher in bullying. He did not appear to have much of a presence in media violence research prior to teaming up with the Zhang lab. What expertise do more recent coauthors have with media violence research? How well do they know their new collaborators' work?

At the end of the day, I think it is safe to say that this is a case of unwanted replication. It is again a stark reminder that peer review is a porous filter. 

One other thing to note. Rodkin passed away in May 2014. The weapons priming paper was first submitted in July of that year. The violent video game study on which Rodkin appears as a coauthor did not get submitted until October 2015. I have no way of knowing what Rodkin's role was on either manuscript, and although I probably will speculate in personal conversations, I won't do so here publicly as that is probably irresponsible. There are certainly ethical ways to handle a situation where a contributor to a research endeavor dies, and I hope that the editors in each instance were made aware of the circumstances at the time of submission.

References:

Tian, J. , Zhang, Q. , Cao, J. and Rodkin, P. (2016). The Short-Term Effect of Online Violent Stimuli on Aggression. Open Journal of Medical Psychology, 5, 35-42. doi: 10.4236/ojmp.2016.52005

Zhang, Q, Tian, J., Cao, J., Zhang, D., & Rodkin, P. (2016). Exposure to weapon pictures and subsequent aggression in adolescence. Personality and Individual Differences, 90, 113-118. doi: 10.1016/j.paid.2015.09.017. 

Thursday, July 9, 2020

Update - what's happened with Zhang Lab papers?

Short answer is that not a lot has happened since the end of last year. More to the point: nothing seems to have happened. I have seen no new English-language publications from the lab. Maybe some specifically Chinese publications have emerged. As of now, I am unaware of them. As of now, there are two retractions (both of which involve reputable scholars, and for whom I can only offer my sympathies), a Corrigendum (which turned out to be only a partial solution - the article needs to be retracted), and a number of errata in what are potentially predatory journals, that are themselves in need of errata. There are a couple journals that have yet to issue so much as a message of concern, despite some glaring errors. If nothing else, the web presence of Qian Zhang at Southwest University in Chongqing has changed considerably over the last year or so. At one point, Zhang had several photos of himself and with eminent US Psychologists from Illinois, along with some statement about SPSS expertise and an enticement to potential grad students to work in his lab. All of that is gone. There is some description of past work, including recent. That's it. Maybe that is progress of a sort. I have no idea of what the CCP has in mind for this particular researcher, nor any particular concern either way. My main concern is that I and my peers can compute accurate effect size estimates, reproduce the findings, and replicate the work. My impression, based on doing a StatCheck scan on a Chinese-language article just prior to Qian Zhang joining the lab at Southwest University, is that the pattern of errors was already in place. This is someone who apparently adapted to a lab culture that was itself in need of improvement. In a toxic academic environment (of which many of us are all too familiar) the convenient way of getting along is the path of least resistance. I wish it were not that way. My guess is that we are looking at a tragedy of errors - one in which there are no villains, just people who made a lot of regrettable choices for the same reason anyone might make regrettable choices in a late capitalist economy. If we get this right in our corner of the sciences, this lab's body of work will be a cautionary tale about the role incentive structures play in terms of career advancement. I suspect there is also a cautionary tale about how eminent psychologists grease the path for success of ambitious researchers, regardless their actual talent and research practices. There is a cautionary tale of coauthors having thorough access to data and codebooks. There is a cautionary tale about editors and peer reviewers having the tools at their disposal to to their jobs as well as possible.There are no happy endings for this particular case. For those of us who really do treasure samples of non-WEIRD populations, I advocate only for making sure that the protocols and data analyses are above board. Else, we get a situation that is truly a mess, and one in which any scholar wishing to extract effect sizes for meta-analyses in the broad area of media violence will be left flustered.

Sunday, October 27, 2019

Erratum to Zhang, Zhang, & Wang (2013) has errors

This is a follow up to my commentary on the following paper:

Zhang, Q. , Zhang, D. & Wang, L. (2013). Is Aggressive Trait Responsible for Violence? Priming Effects of Aggressive Words and Violent Movies. Psychology, 4, 96-100. doi: 10.4236/psych.2013.42013

The erratum can be found here.

It is disheartening when an erratum ends up being more problematic than the original published article. One thing that struck me immediately is that the authors continue to insist that they ran a MANCOVA. As I stated previously:
It is unclear just how a MANCOVA would be appropriate as the only DV that the authors consider for the remaining analyses is a difference score. MANOVA and MANCOVA are appropriate analytic techniques for situations in which multiple DVs are analyzed simultaneously. The authors fail to list a covariate. Maybe it is gender? Hard to say. Without an adequate explanation, we as readers are left to guess. Even if a MANCOVA were appropriate, Table 4 is a case study in how not to set up a MANCOVA table. Authors should be explicit about what they are doing as possible. I can read Method and Results sections just fine, thank you. I cannot, however, read minds.

In essence, my initial complaint remains unaddressed.  One change, Table 4 is now Table 1, and it has different numbers in it. Great. I still have no idea (nor would any reasonably-minded reader), based on the description given, what the authors used as a covariate nor do I know what purported multiple DVs were used simultaneously. This is not an analysis I use very often in my own work, although I have certainly done so in the past. I do have an idea of how MANOVA and MANCOVA tables would be set up, and how those analyses would be described. I did a fair amount of that for my first year project at Mizzou a long time ago. The authors used as their DV a difference score (diff between RT aggressive words vs RT nonaggressive words), which would rule out the need for a MANOVA. And since no covariate is specified, a MANCOVA would be ruled out. I am going to make a wild guess that the partial summary table that comprises Table 1 will end up being nonsensical as have been similar tables generated in papers by this lab, including errata and corrigenda. I don't expect to be able to generate the necessary error MS, which I could then use to estimate the pooled SD.

I also want to note that the description of Table 2 as characterized by the authors and the numbers in Table 2 do not match up. I find that troubling. I am assuming that the authors mislabeled the columns, and intended for the low trait and high trait columns to be reversed. It is still sloppy.
At least when I ran this document through Statcheck, the findings, as reported, appeared clean - no inconsistencies and no decision inconsistencies. I wish that provided cold comfort. Since I don't know if I can trust any of what I have read in either the original document or the current erratum, I am not sure that I there is any comfort to be had.

What saddens me is that so much media violence research is based on WEIRD samples. That influences the generalizability of the findings. That also limits the scope of any skepticism I and my peers might have about media violence effects. We need good non-WEIRD research. So the fact that there is a lab that is generating a lot of research that is non-WEIRD, but is riddled with errors is a major disappointment.

At this juncture, the only cold comfort I would find is if the lot of the problematic studies from this lab were retracted. I do not say that lightly. I view retraction as a last resort, when there is no reasonable way for the record to be corrected without removing the paper itself. Doing so appears to be necessary for at least a few reasons. One, meta-analysts might try to use this research - either the original article or the erratum (or both if they are not paying attention) to generate effect size estimates. If we cannot trust the effect size estimates we generate, it's pretty much game over. Two, given that in a globalized market we all consume much of the same media (or at least the same genres), it makes sense to have evidence from not only WEIRD samples but also non-WEIRD samples. Some of us might try to understand just how violent media affect samples from non-WEIRD populations in order to understand if our understanding of these phenomena are universal. The findings generated from this paper and from this lab more broadly do not contribute to that understanding. If anything, the findings detract from our ability to get any closer to the truth. Three, the general public latches on to whatever seems real. If the findings are bogus - either due to gross incompetence or fraud - then the public is essentially being fleeced, which to me is simply unacceptable. The Chinese taxpayers deserved better. So do all of us who are global citizens.


Monday, October 14, 2019

"It's beginning to, and back again": Zheng and Zhang (2016) part 2

A few months ago, I blogged about an article by Zheng and Zhang (2016) that appeared in Social Behavior and Personality. I thought it would be useful to briefly return to this particular article as it was (unbeknownst to me at the time) my first exposure to that lab's work, and because I think it might be helpful if you all are seeing what I am seeing when I read the article.

I don't think I need to re-litigate my reasoning for recommending a rejection when I was a peer reviewer on that particular article, nor my disappointment that the article still was published anyway. Water under the bridge. I think what I want to do is to share some screen shots of the analyses in question as well as to note a few other odds and ends that always bugged me about that particular paper.

I am keeping my focus to Study 2, as that seems to be the portion of the paper that is most problematic. Keep in mind that there were 240 children who participated in the experiment. One of the burning questions is why the degrees of freedom in the denominator for so many of the analyses were so low. As the authors provided no descriptive statistics (including n's) it is often difficult to know exactly what is happening, but I might have a guess. If you follow the Zhang lab's progression since near the start of this decade, sample sizes have increased in their published work. I have a sneaking hunch that the authors copied and pasted text from prior articles and did not necessarily adequately update the degrees of freedom reported. The df for simple effects analyses may actually be correct, but there is no real way of knowing given the lack of descriptive statistics reported.

One problem is that there seemed to be something of a shifting dependent variable (DV). In the first analysis where the authors attempted to establish a main effect, the authors only used the mean reaction times (rt) for aggressive words as the DV. In subsequent analyses, the authors used a mean difference in reaction times (rt neutral minus rt aggressive) as the DV. That created some confusion already.

So let's start with the main analysis, as I do have a screen shot I used in a tweet a while back:

So you see the problem I am seeing already. The analysis itself is nonsensical. There is no way to say that violent video games primed aggressive thoughts in children who played the ostensibly violent game as there was no basis for comparison (i.e, rt for non-aggressive words). There is a reason why I and my colleagues in the Mizzou Aggression Lab a couple decades ago computed a difference score between rts for aggressive words and rts for non-aggressive words and used it as our DV when we ran pronunciation tasks, lexical decision tasks, etc. If the authors were not willing to do that much, then a complex between/within ANOVA in which the interaction term would have been the main focus would have been appropriate. Okay. Enough of that. Make note of the degrees of freedom (df) in the denominator. With 240 participants, there is no way one is going to end up with 68 df in the denominator for a main effects analysis.

Let's look at the rest of the results. The Game Type by Gender interaction analyses were, um, a bit unusual.

First let's let it soak in that the authors claim to be running a four-way ANOVA, but there appear to be only three independent variables: game type, gender, and trait aggressiveness. Where is the fourth variable hiding? Something is already amiss. Now note that first analysis goes back to main effects. Here the difference between rts for aggressive and nonaggressive words is used as the DV, unlike the prior analysis that only examined rts for aggressive words as the DV. Bad news, though: if we believe the df as reported, a Statcheck analysis shows that the reported F could not be significant. Bummer. Statcheck also found that although the reported F for the game type by gender interaction - F(1,62) = 4.89 - was significant, it was at the p = .031. Note again, that is assuming the df can be taken at face value. The authors do report the mean difference scores for boys playing violent games and nonviolent games, but do not do so for girls in either condition. I found the lack of descriptive statistical data to be vexing, to say the least.

How about the analyses examining a potential interaction of game type and trait aggressiveness? That doesn't exactly look great:

Although Statcheck reports no obvious decision errors for the primary interaction effect or the simple effects, the df reported are, for lack of a better way to phrase this, all over the place. The lack of descriptive statistics makes it difficult to diagnose exactly what is going on. Then the authors go on to report a 3-way interaction as non-significant, when a Statcheck analysis indicates that it would be. If there were a significant 3-way interaction, that would require some considerable effort to carefully characterize the interaction, and to carefully graphically portray the interaction.

It also helps to go back and look at the Method section and see how the authors determined how many trials each participant would experience in the experiment:


As I stated previously:

The authors selected 60 goal words for their reaction time task: 30 aggressive and 30 non-aggressive. These goal words are presented individually in four blocks of trials. The authors claim that their participants completed 120 trials total, when the actual total would appear to be 240 trials. I had fewer trials for adult participants in an experiment I ran over a couple decades ago and that was a nearly hour-long ordeal for my participants. I can only imagine the heroic level of attention and perseverance required of these children to complete this particular experiment. I do have to wonder if the authors tested for potential fatigue or practice effects that might have been detectable across blocks of trials. Doing so was standard operating procedure in our lab in the Aggression Lab at Mizzou back in the 1990s. Reporting those findings would have also been done - at least in a footnote when submitted for publication.

Finally, I just want to say something about the way the authors described the personality measure they used. The authors appeared to be interested in obtaining an overall assessment of aggressiveness. The Buss & Perry AQ is arguably defensible for such an endeavor. The authors have a tendency to repeat the original reliability coefficients reported by Buss and Perry (1992), but given that the authors only examined overall trait aggressiveness, and given that they presumably had to translate this instrument into Chinese, the authors would have been better served by reporting the reliability coefficient(s) that they specifically obtained, rather than doing little more than copying and pasting the same basic statement they make in other papers published by the Zhang lab. It really takes getting to the General Discussion section before the authors even obliquely mention that this instrument was translated, as well as to more specifically recommend an adaptation of the BPAQ specifically for Chinese-speaking and reading individuals.

This was a paper that had so many question marks that it should never have been published in the first place. That it did somehow slip through the peer review system is indeed unfortunate. If the authors are unable or unwilling to make the necessary corrections, it is up to the editorial team at the journal to do so. I hope that they will in due time. I know that I have asked.

If not retracted, any corrections would have to report the necessary descriptive statistics upon which the analyses for Study 2 were based, as well as provide the correct inferential statistics: accurate F-tests, df, and p-values. Yes, that means tables would be necessary. That is not a bad thing. The actual Coefficient Alphas used in the specific study for their version of the BPAQ should be reported, instead of simply repeating what Buss and Perry reported for the original English language version of the instrument in the previous century. The editorial team should insist on examining the original data themselves so that they can confirm that any corrections made are indeed correct, or so that they can determine that the data set is so hopelessly botched that the findings reported cannot be trusted, hence necessitating a retraction.

How all this landed on my radar is really just the luck of the draw. I was asked to review a paper in 2014, and I had the time and interest in doing so. The topic of the paper was in my wheelhouse, so I agreed to do so. I recommended a rejection, which in hindsight was sound. I moved on with my life. A couple years later I would read a weapons priming effect paper that was really odd and with reported analyses that were difficult to trust. I didn't make the connection until an ex-coauthor of mine appeared on a paper that appeared to have originated from this lab. At that point I scoured the databases until I could locate every English-language paper published by this lab, and discovered that this specific paper - which I recommended rejecting - had been published as well. In the process, I was able to notice that there was a distinct similarity among all the papers - how they were formatted, the types of analyses, and the types of data analytic errors. I realized pretty quickly that "holy forking shirtballs, this is awful." I honestly don't know if what I have read in this series of papers amounts to gross incompetence or fraud. I do know that it does not belong in the published record.

To be continued (unfortunately)....

Reference:

Zheng, J., & Zhang, Q. (2016). Priming effect of computer game violence on children’s aggression levels. Social Behavior and Personality: An International Journal, 44(10), 1747–1759. doi:10.2224/sbp.2016.44.10.1747

Footnote: The lyric comes from the chorus in "German Shepherds" by Wire. Toward the the end of this post near the end of the last paragraph, I make a reference to some common expressions used in the TV series, The Good Place.

Saturday, August 6, 2022

The struggle continues: Zheng and Zhang (2016) Pt. 4

Let's focus on Study 1 of Zheng and Zhang (2016). It should have been fairly simple, at least in terms of data reporting. However it happened, the authors honed into two video games that they thought were equivalent in terms of any confound aside from violent content, and merely needed to run a pilot study to demonstrate that they could back that up with solid evidence. It should have been a slam-dunk.

Not so fast.

The good news is that, unlike Study 2, the analyses of the age data are actually mathematically plausible. That's swell. I noticed that the authors had some Likert-style questions to rate the games on a variety of dimensions, which makes sense. The scaling was reported to be on a 1 to 5 scale, in which 1 meant very low and 5 meant very high for each dimension. My intention was to focus on Tables 1 and 2. If a 1 to 5 Likert scale was used for each of the items used to rate the games, there were some problems. One glaring problem is that there is no way that there could be means above 5. And yet, for Violent Content and Violent Images dimensions, the mean was definitely above 5 in each case. That does not compute. I have no idea what scaling was used on the questionnaires actually used. I can perhaps assume a 1 to 7 Likert scale. Certainly doing so would make some means and standard deviations that seemed mathematically impossible seem at least with in the realm of plausibility. But there is no way to know. We do not have the data. We do not have any of the materials and protocols. We have to take everything on faith. I had intended to have a set of images of SPRITE analyses on Table 1 and Table 2, but didn't see the point. 

Then we have the usual problem with degrees of freedom. With a 2x2 mixed ANOVA, with game type as a repeated measure and "gender" as a between-subjects factor, the degrees of freedom would not have deviated much from the sample size of 220. I think we can all agree with that. Degrees of freedom below 100 would be impossible. And yet the analyses reported do just that. It does not help much that Table 1 is mislabeled as t-test results. If we assumed paired sample t-tests, degrees of freedom for each item would have been 219. Again, the reported degrees of freedom do not compute.

What I can say with some certainty is that Zheng and Zhang (2016) should not be included in any meta-analysis addressing violent video games and aggression or media violence and aggression. My efforts to address some of these issues with the editorial staff never went very far. It's so funny how problems with a published paper lead editorial staff to go on vacation. I get it. I'd rather be out of town and away from email contact when someone emails (with evidence) concerns about a published paper. Unfortunately, if the data and analyses cannot be trusted, we have a problem. This is precisely the sort of paper that, once published, ends up included in meta-analyses. Meta-analysts who would rather exclude findings that are, at best, questionable will be pressured to include such papers anyway. How much that biases the overall findings is clearly a concern. And yet the attitude seems to be to let it go. The attitude is that the status quo is sufficient. One flawed study surely could not hurt that much? We simply don't know. The same lab persisted, with samples of over 3,000, to publish research relevant to media violence researchers. Several of those papers ended up retracted. Others probably should have been, but probably won't due to whatever political reasons one might imagine. 

All I can say is the truth is there. I've tried to lay it out. If someone wants to run with it and help make our science a bit better, I welcome you and your efforts.

Friday, December 6, 2019

For those visiting from Retraction Watch:

Retraction Watch posted an article about two retractions of articles in which Qian Zhang of Southwest University in China was the lead author. Since some of you might be interested in what I've documented about other published articles from Zhang's lab, your best bet is to either type Zhang in the search field for this blog. Or just follow this link, where I have done the work for you. I'll have more to say about these specific articles in a little bit. I think I documented some of my concerns on Twitter last year and pseudonymously on PubPeer. In the meantime, I am relieved to see two very flawed articles removed from the published record. Joe Hilgard deserves a tremendous amount of credit for his work reanalyzing some data he was able to obtain from the lab (and his meticulous documentation of the flaws in these papers), and for his persistence in contacting the Editor in Chief of Youth and Society. I am also grateful for tools like Statcheck, which enabled me to very quickly spot some of the problems with these papers.

Wednesday, October 7, 2020

The Zhang Lab rides again

If you've read this blog long enough, you're familiar with the work of Qian Zhang of Southwest University in China. You are already well aware that there are some serious problems with many of the papers he has co-authored (either as a first author or a more secondary co-author) over the years. His more recent papers have been on the surface of higher quality, but it sometimes doesn't take much to realize that there are still substantial problems. Bottom line is that if you see his name mentioned here, it's not good news.

Case in point: Dr. Zhang has a new paper out that purports to examine the link between viewing prosocial cartoons and a reduction in aggressive cognition and behavior. I was alerted to this paper by Joe Hilgard. On the surface, a simple Statcheck run looked good. Initially I lamented the lack of tangible data to reproduce the analyses. Dr. Hilgard pointed me to where the data were stored (which kudos to this lab for doing so). I ported the dataset into my current version of jamovi and successfully reproduced the analyses reported in the paper. So far so good. Then I had that sinking realization something was still wrong. The data set only contained data for reaction time data for weapon images (which the authors use for the DV in the paper and analyses - a fact Dr. Hilgard had already arrived at before I did my work here). However, the authors should also had data on reaction times for neutral images that were not included in the data set. The appropriate DV would have been a difference score between reaction times for aggressive images (in this case, weapons) and reaction times neutral images. That difference score would be the proper measure of accessibility of aggressive cognitions. 

As of this writing, the last author on the paper had been contacted, and I trust this last author to do the right thing here. At bare minimum, a reanalysis needs to be conducted in order to ascertain that prosocial cartoons really did lead to a decrease in the relative accessibility of aggressive cognition. As of now, the paper cannot adequately address that claim. There is this funny gray area between what we consider published and in press. The paper has already been accepted, and some version of it has been made available online. My hope is that the last author, an American researcher with a solid reputation, is able to get the matter resolved satisfactorily, however that turns out. Maybe a simple correction suffices. It is possible that a properly calculated cognitive DV yields the same basic findings as the incorrect cognitive DV. If it does not, then many of the conclusions of the paper may need to be rethought and rewritten. If so, the topic is of enough theoretical and practical interest that perhaps a sympathetic editor and publisher will still be okay with a corrigendum, regardless of how the ultimate findings flush out. If a retraction is necessary at this stage, it would be far less painful than after it is already officially in print. This is a matter of making sure that those of us who might still be tempted to conduct meta-analyses in this broad area of media violence have the correct findings when estimating effect sizes, that those who might be using this literature to advocate for policy changes have the right information before coming across as grossly uninformed. For the good of the order, I hope this matter is taken care of quickly. In the meantime, I'd warn against citing this particular paper unless and until at least some sort of correction has been published. 

Reference: https://doi.org/10.1016/j.childyouth.2020.105498