Matthias J. Becker on the Launch of Digital Hate Review
Justin Hendrix / Sep 20, 2026Audio of this conversation is available via your favorite podcast service.
The study of how hate is expressed in digital media and its relationship to offline harassment and violence is a growing field that touches public policy, content moderation and platform policy, AI governance, and a variety of other topics at the intersection of technology and public discourse. This week, a new journal called the Digital Hate Review launched to give a new home to research and discourse on these issues. To learn more, I spoke to its editor-in-chief, Matthias J. Becker, who has spent 15 years tracking how hate speech evolves online.
What follows is a lightly edited transcript of the discussion.
Matthias J. Becker:
So my name is Mathias Jacob Becker. I'm a cognitive linguist and discourse researcher. I am the AddressHate Research Scholar at NYU's Center for the Study of Antisemitism, where I'm also leading a project called Decoding Hate, and I work very closely with the think tank AddressHate. And here today I want to speak as the editor-in-chief of Digital Hate Review, the new open access journal published by AddressHate, whose first issue launched last Monday.
Justin Hendrix:
And I'm excited to talk to you about this subject and also about the launch of the Digital Hate Review. We're going to get into that, the motivation for that, and some of the issues and ideas that you're addressing, where you hope this review will go in the future. My listeners have probably become familiar with something I often like to do when I feel like I'm introducing someone new to them, which is to kind of ask about their intellectual journey, their intellectual curiosities. So how did you go from cognitive linguistics to working on antisemitism, to thinking about the media, to thinking about how antisemitism, misogyny, racism express themselves online? Explain that kind of trajectory for us.
Matthias J. Becker:
Yeah, the pleasure. And thank you for having me. It's a pleasure to be here. So my background, as I said, is cognitive linguistics. I was always very interested in languages and the relationship between languages or language and power. This whole relationship is, I think is something extremely interesting when you look at the history of discourse, the history of propaganda, the history of how rhetoric has been used in history and of course also today in order to, for instance, justify violence, in order to promote certain hate ideologies, conspiracy narratives. So these things are very closely connected. I really started with philology, philosophy in Berlin. I did my PhD at TU Berlin on antisemitism, specifically on antisemitism, but also more generally on hate ideologies in language, in everyday language online. So for me, it was not really the point to study language just for the sake of learning new words and study vocabulary and the repertoire of a different language, but actually to see how different actors online in particular use language in order to frame incidents, ideas in a specific way, because words can tell you a lot about the attitude, the thinking, the feelings of the speaker.
And over time, it was just really a tiny step towards antisemitism because it really became quite a thing 10, 15 years ago already in Berlin and in Germany in general that you could feel this return of hateful ideas about Jews even in mainstream discourse. So my starting point was actually racism and nationalism, and then I became more more interested in antisemitism and how it operates in a country like Germany with all its past. And I did my PhD on that particular topic and in 2020 I started, well, probably one of the biggest projects on hate speech online, Decoding Antisemitism at TU Berlin funded by the Alfred Landecker Foundation in close collaboration with King's College in London and HTW Berlin. And with a team of around 25 scholars from different disciplines, we were looking at antisemitic rhetoric in mainstream online environments and using an interdisciplinary approach starting with the fine-grained granular linguistic analysis of the content we could see in these mainstream online contexts and trying to bridge the gap between this qualitative analysis approach with AI-based approaches at scale.
And the product was, or the result was a multilingual corpus for English, French, and German of more than 300,000 annotated comments and an open access lexicon called decoding antisemitism that till today has been accessed nearly half a million times. So I could see in all these years that there's a tremendous interest in the phenomenon of digital hate, that there's a lot of uncertainty of how to assess that. Everybody knows that there are also a lot of debates about the right definition when it comes to hate speech, when it comes to antisemitism, but also other hate ideologies. And I wanted to bridge the gap between this kind of theoretical assessment, the historical analysis of hate and exclusion, how it evolved from a diachronic perspective, related to really applied linguistics, corpus linguistics, and of course bridge the gap to LLM enhanced methods because this is just really an important combination in order to say something substantial about the current phenomena and trends on different social media platforms.
Justin Hendrix:
So I want to get a little bit more into your research and of course your intent with digital hate review and the community you're trying to build around that. But before we do move on from where you brought us to there on the state of the research on the phenomenon of hate, hate speech, online hate, where are we right now? I mean obviously I've been following these things for a while. It feels like there's been an enormous amount of effort, enormous amount of computational social science that's gone in to answer these questions. And unfortunately, as you say, a return, in some cases, maybe that's not enough of a word to describe the trajectory of things. I happen to be in Berlin talking to you today, and there's an election here this weekend, a local one. And so there are political signs all around that it's been interesting to me to see the amount of signage and enthusiasm for the AfD, for instance.
What do you think this field is up to right now? Where has it got to? How would you describe it in terms of just its health or its capacity?
Matthias J. Becker:
This is a very good question. I think looking at how the field has evolved in the last. I mean, I started working in the field 15 years ago and I was immediately interested in verbal violence and digital hate on different social media contexts. I would say that back in the day, 15 years ago, digital hate studies was more kind of a niche concern. So there were case studies of extremist forums. There were attempts to basically count slurs. There was single platform snapshots. And I think a lot of things have changed because we just have a richer diversity of research designs and tools that we can use. I think first of all, a lot of research centers and also activists and civil society organizations understand that it's not always easy to identify hate because it's not just the slur, it's not just incitement to violence, any form of hate and exclusion, especially when it arrives in politically moderate contexts, and this is something that we can see globally today.
We can see that there's a whole diversity of implicit hate speech. There are codes, illusions, indirect speech like rhetorical questions, people use irony and jokes. And of course you have the whole layer of text image relations, audiovisual content that all that makes it so much harder to measure the current state of affairs and the developments and trends that we need to have on our radar. So I think this is something that has happened in the field that people are more aware of the diversity and complexity of the research object so that you need to have context, you need to have world knowledge, historical knowledge, you need to know how a different online community communicate. I don't know a 30-year-old web user on X will communicate hateful ideas against women in a different way than a 15-year-old on TikTok. So you have a diversification of rhetoric of the retorted devices that people use, and therefore you need to have another form of assessing these things.
And of course the scale changed. So we moved from a rather anecdotal assessment of online hate to annotated data sets, shared data sets, a systematic, systematic taxonomies. As I said, we had a multilingual corpus of more than 300,000 manually annotated comments. And I would also say that the scope changed massively. So back in the day, there were a lot of research projects on single platforms and now you see more and more research on cross-platform dynamics from isolated content to discourse events, how a news event sends waves of hate through comment sections within hours. And I think actually these new opportunities that we have, again, led to a much more diverse or much more sophisticated understanding of the problem. Implicitness dominates mainstream spaces. Explicit hate has been pushed to margins. Coded language circulates broadly. The context is highly relevant. We need to look at the trigger events, structure the phenomenon.
And of course we also have more and more research on the correlations between online hate and offline violence. There's a range of case studies that actually look at, for instance, anti-refugee sentiment on certain platforms and how they predict attacks on refugees in different, for instance, European contexts. And of course this correlation is something that we need to look more in all details because it is something that will happen on a more and more regular basis. So the tone that you can find online will of course have an impact on the behavior in offline contexts. I think something that is contested in this context, okay, there's a lot of agreement. It's not simple to define hate speech, digital hate, where to draw the line between the different concepts, the different narratives that for instance, carry a conspiracy narrative. But there is of course also a lot of room for debate when it comes to the prevalence.
How much hate there is depends heavily on definitions and sampling. Again, I spoke about the effects for correlations. The correlation to find correlations I think is quite easy. It's not always easy to identify causation. So this is another approach. And of course there's a lot of discussion about the right form of interventions. So de-platforming works conditionally. And there are researchers that found that de-platforming reduced activity and toxicity around figures like Alex Jones, but of course this is
Just reducing the reach and displacement, even though they're both real, there are still certain forms that as soon as you make it impossible to communicate in a certain sphere, the communication can continue just on other platforms. And sometimes it's even more and more difficult to track because the platform, the new platform makes it more difficult to analyze the patterns that we can find there. And just for your last question about computational social science, of course this had a massive impact on the field. So it gave us scale. We have longitudinal designs. We have corporate data sets of millions of comments that we can read. But I would say the decisive progress came from hybrid designs. So you have computational scale grounded in qualitative context sensitive annotation because for instance, implicit hate cannot be read off surface features that has levels for founding of my work as well.
And the first DHR issue, and we can talk about this in a minute, shows where that hybrid frontier stands today. So there are colleagues who will talk about this relationship between these two sides, a very fruitful conversation between two research or research approaches. But nevertheless, computational methods certainly transformed what we can see. So the interpretive work, understanding what we are looking at is as necessary as before. The field has become much, much larger, more empirical and more computational, but it has not become conceptually settled, I would argue.
Justin Hendrix:
Well, let me ask a question about impact. And I know you yourself have worked with in past the State Department, the European Union, I believe at least one of the major platforms. In years past, of course, in the kind of heyday of trust and safety, for instance, there was an enormous amount of effort around questions around hate speech, questions around various other phenomena that would generally fall into the kind of categories we're discussing here. I'm wondering what your view on that side of the field is. Of course, things have changed, policies have changed, staffing has changed, governments have changed, priorities have changed. What's your view on the kind of, I suppose, demand side for your research? What are you seeing in terms of the type of uptake for the types of ideas you're now able to generate with these methods? How is it potentially being consumed or even advanced into change?
Matthias J. Becker:
This is a very good question. And of course the impact side is absolutely an essential part. I wouldn't even know where to start. There's so many domains, so many potential target audiences that we can reach out to with this kind of work because it doesn't matter if you talk to a university professor or a teacher at a K through 12 schools, they are all really, really interested in knowing how the debate culture online looks like. And not just through exploratory anecdotal studies that very often cannot really present results that are reliable because very often you can actually find anything online. It's a question of how representative your data sets and your results eventually are. But in their everyday work, in their work of, for instance, informing young people, students of different age groups and different age groups, how hatred looks like, what kind of conspiracy narrative we can find, how anti-democratic resentment looks like, they of course need to know in the first place how these patterns occur.
What are the main narratives, the main, what's the word choice? What are the argumentative patterns that you can find online that hijack a lot of especially younger generations into a belief system which is quite authoritarian, which is quite hateful, that is exclusionary. So in order to build something that is addressing these issues, you need, of course, in the first place, you need to understand where we currently stand. If you want to build curricula, you need to know what your target group, your students are potentially confronted with. The same was true with, of course, policymakers, also those audiences that actually have a dialogue with social media companies. They need to know where we currently stand in order to formulate clearer arguments, also to strengthen the discussion and also raise the likelihood that we can actually make a difference. If we are just quoting some numbers, but we can't really prove it based on reliable results, based on visualizations, based on dashboards, like where do we stand when it comes to anti-Muslim hate, where do we stand regarding antisemitism, misogyny, queerphobia, you name it, then I think the debate cannot really move forward.
So I think there it would be quite interesting to see how this can improve our current situation because we have, again, not much knowledge about online discourse. We have, again, a range of approaches, a range of case studies, but it's very, very difficult to connect them because each of those have their own research designs, their own categories, their own methods. So that makes it really, really hard. And as soon as we actually find a more comprehensive approach through an interdisciplinary format, we can make a difference on the level of education policy that I just mentioned also regarding law. There's a lot of interest in the legal domain to understand how hate speech, hate communication, also when we talk about how this currently looks like, asking the question of how much our current definitions that are in place suffice for really a moving target. So long story short, I think my idea would be to look at social media to apply a comprehensive form of social media studies of digital hate studies in order to inform different target audiences so that they can act in their own domain in real time or just in very, very short intervals.
Because at the moment we have a couple of reports that come out on a regular basis. This is wonderful work. I don't want to trivialize that in any way, but very often these reports are already outdated when they get published. I think we need to find a better way for this kind of knowledge transfer from digital hate studies, social media studies in general to these different domains. And there's a strong appetite for that. There's a lot of, just because of the work that I've been doing at NYU and what I can see when it comes to AddressHate of the nonprofit organization that I advise, there's just strong interest from different corners of the field. And here I think there's a lot of room for improvement for building an ecosystem that allows us to understand society in a much better way than ever before.
Justin Hendrix:
And I assume that goes a long way to explaining the motivation behind the digital hate review.
Matthias J. Becker:
Yeah, that's exactly right. In general, something that I haven't mentioned before, like my whole trajectory in Berlin and in Europe, then New York from November last year onwards, there's a clear move on my end to broaden the comparative framework that I developed back then in 2020 with decoding antisemitism. The new project that I've been leading is called Decoding Hate, where we include anti-Black racism, misogyny, anti-Asian racism, and potentially in the near future also other hate ideologies. This is something I've been working on also closely with AddressHate as mentioned before. And here, without intending to equate these hate ideologies because each hate ideology has its own conceptual repertoire, its own stereotypes, concepts, narratives, and so on and so forth. But just looking at them from a bird's eye view and see, for instance, what incidents in the real world, in the offline world trigger what kind of antisemitic or anti-Asian conspiracy narratives, what kind of forms of dehumanization, demonization we can find on the side of anti-Black racism and misogyny.
All this will just help us understand or assess the current state of affairs in a much better way. And this is basically also what I wanted to build with digital hate review with the DHR. So as soon as AddressHate and Joshua Laterman, the founder, and Emily Carmeli, the president of AddressHate, and I had a conversation about this nearly a year ago, we actually came up with the idea of really building an academic journal that connects the dots between these different work fields. So our idea was really to bring these different dispersed disciplines together because they all often use different concepts, methods, publication cultures, there are different timescales. So a computational scientist developing a detection system, for instance, or a linguist like me interested in implicit hate or a political scientist examining polarization. Very often these diverse research communities or milieus don't really talk to each other and they could learn a lot from each other, I think.
Just looking at the field of antisemitism studies, I think there are lots of overlaps with racism studies and against studies of other hate ideologies. And this is exactly what DHR tries or intends to build, some kind of a venue, a hub that also allows future authors, contributors to play around with different tools for different software that we have on our website. So authors can also use that software. They can be in touch with each other. They can actually have a better dialogue about how we can conduct hate studies, how we can conduct social network analysis, what kind of tools can be used in order to fill the gaps, the conceptual gaps in our research design. And I think this whole fragmentation that we can see until today matters a lot because digital hate is not, of course, a marginal problem. It touches some of the major challenges democratic societies are dealing with.
We have polarization, we have the erosion of trust in democratic institutions. The quality of public discourse is really deteriorating the issue of safety, the safety issues of minority communities, the omnipresence of political violence, and of course, increasingly the question of governance of AI systems that can produce and amplify harmful content at enormous scale. So none of these problems belongs to a single discipline. If you study them in disciplinary silos and isolation, we risk understanding individual mechanisms without understanding the larger system in which they interact. And that for me is the deeper rationale for the journal. So in here, in the inaugural editorial, I put it this way, that the first issue is both a beginning and also an argument at the same time. The argument is that hate studies in digital space has reached enough conceptual methodological, also empirical maturity to justify a dedicated home while still being fragmented enough to need one.
So the field is producing more research than ever, but not enough conversation. And that's the problem the journal wants to address here.
Justin Hendrix:
Well, I want to give the listener a little bit of a sense of what they might encounter. And of course, they can go themselves to digitalhatereview.com and check out the latest articles, the first edition that's up. But just back to the point of artificial intelligence, you addressed that in at least one of the articles that's up, this paper by, hopefully I'll get this pronunciation correct, you can correct me if I don't, this paper by Mykola Makhortykh and Elizaveta Kuznetsova, “Routine Distortion,” this idea of Holocaust distortion as a kind of byproduct of generative AI architecture. I'm not going to ask you to necessarily fully represent these researchers' work, but maybe just for the sake of giving a listener a sense of what they'll encounter at Digital Hate Review, what's this about?
Matthias J. Becker:
Yeah, so we have a range of really, really interesting research pieces. Before we get into this, I just wanted to highlight one thing, and I think this makes the DHR quite unique. As I mentioned before, it's not just research in an academic journal, but actually a hub that tries to connect different environments, different research fields. That's of course quite obvious and also really needed, but there's also an internal structure that aims at reaching out to different target audiences. So as I said before, we have of course a range of research articles in the research section that provides a home for empirical and analytical scholarship across the disciplines that constitute hate studies. But we also have a legal forum that is designed as an asynchronous dialogue between legal experts on questions connected to digital hate, platform governance, AI, freedom of expression, accountability, and so on and so forth.
And the idea here was rather than publishing isolated legal essays, it is organized as a kind of a sequence. An opening essay starts and sets out a legal problem or position, and then subsequent contributors respond from their own comparative jurisdictional perspectives. So those responses can in turn be answered and the discussion accumulated as a citable intellectual record. So we have really a kind of a dialogue, like a chat between different legal experts about questions that should be relevant for all of us. So we also asked our contributors to use a language that is accessible, but of course with the right rigor and the precision that they should apply in their work. But the most important part is actually to see where do we currently stand? What is the question about liability issues? What about the debate culture currently about freedom of speech in times when the society discourse becomes more and more polarized and the probability of hateful eruptions is just increasing?
And so this is the legal forum, and then we have a third section that's called Perspectives. And here we invite trust and safety teams, platform engineers, civil society organizations, researchers, policymakers to actually add their contribution and to tell from their perspective where we currently stand. So kind of a three layered approach, starting with research articles, getting to the legal forum, which is clearly designed as an exchange forum between these different sites in that particular environment and perspectives that kind of bridges the gap between the world of research and civil society and other domains. And just for the research articles, for the first issue, we could actually bring a range of scholars on board. She did really fascinating work. The work you just mentioned by Makhortykh and Kuznetsova is called “Routine Distortion.” And the idea here is to get into the topic of generative AI and historical memory, but not by focusing on, I would call it spectacular fabrication.
So it's not about deep fakes that kind of distort the memory of the genocide committed by the Germans or just rearranging facts and it becomes a big headline or the whole debate about Grok and what information of Nazism you can find, but it's something that they call routine distortion. And their intervention begins by distinguishing between viral distortion from what they call this kind of routine distortion, which is different. It is basically unprompted misrepresentation of historical facts that emerges during completely ordinary interaction with generative AI. So we did a range of testing and that matters because the underlying systems remain unreliable for precisely this kind of historical inquiry. So the audit work they review found historically accurate answers in only roughly half of responses in some studies. And even under the best conditions, only around 70% of prompts were answered correctly. And the errors are not merely wrong dates.
The systems generated non-existent war crime trials. They invented Memory laws, they fabricated quotations from witnesses and perpetrators and attributed responsibility for atrocities in an incorrect way. So this whole observation, their intervention showed us that this goes beyond hallucination. Response can contain no invented fact that stands out that through flattening, for instance, complexity, through a selective framing, inconsistent or misleading contextualization, this is how the society discourse or the collective memory about the Holocaust, and of course potentially also of other historical atrocities can shift and can move in a different direction. Of course, that's really important because search is gradually moving from giving users sources to giving users answers. So if the interface gives you an answer and you never visit the museum or the archive or historical source behind it, generative AI becomes a new historical gatekeeper. This definitely has, therefore, a massive effect. This is not only a content problem, it is a change in the infrastructure through which societies encounter the past.
Justin Hendrix:
That's fascinating. And we've had a couple pieces on tech policy press that have been dealt with questions around this type of phenomenon, but to see it kind of play out in this way I think is both really interesting, but also incredibly concerning as we see these technologies rolled out across the board. And you even think about some of the debates we're having over the role of generative AI in schools. And these are the types of questions I feel like we desperately need to answer before we make these tools available to students to try to use to understand the world until we understand these types of biases and other types of phenomena.
Matthias J. Becker:
And the authors are basically also making it very clear, they basically ask the field to stop concentrating only on, well, spectacular AI failures and start auditing ordinary information seeking at population scale. So they propose in their article practical directions that we need expert curated knowledge resources, we need regular multilingual audit and targeted AI literacy. And their point is not that every possible error can be eliminated, but that as soon as you have a high frequency of specific historical questions, that you could make this whole scenario much, much safer through systematic intervention. And that's also something that I really like about that contribution, that it's not only looking into a challenging problem that I guess a lot of people have heard about, but also just laying out what could be done and how this could look like in a way that is actually manageable because the amount of work and just the sheer diversity of answers that you can get, and of course something that is already part of a podcast scene for a couple of years, how do we deal with chatbots that kind of rewrite history and are able, are powerful enough to produce seemingly historical sources that people actually believe in on a massive scale is just a challenge that is not, in my opinion, not present enough in a political debate culture.
Justin Hendrix:
Well, another piece that I wanted to ask about, which is kind of the flip side in a way, there's an enormous amount of focus right now on using large language models for content moderation. Of course, using machine learning in content moderation has been something that has been going on for at least a couple of decades at scale. But this kind of question about what to do with our language and how it changes and the extent to which the terminologies that people are using in some cases to other people to engage in hate, to engage in various other forms of harassment and other types of dangerous speech can be difficult for machines to adapt. So this piece you have from Patrick Wu and his co-authors on a computational approach discovering emerging hate speech, I think this is interesting and probably worth a read for anybody out there who's thinking about trying to automate content moderation on a platform.
Matthias J. Becker:
Yes, certainly. I mean this is one of the major problems. Most hate speech detection would ask the question, does this piece of text contain hate? And these authors here ask an earlier question, how do we discover the vocabulary of hate before our detection systems even know that this vocabulary exists? So they could see that there's a clear asymmetry in content moderation because hate vocabulary evolves over time. Detection systems are usually trained on yesterday's language. But the problem is of course that online communities invent euphemisms. You have altered spellings as well. You have coded expressions, you have apparently innocuous substitutes precisely because existing vocabulary is detectable. I wouldn't be limited to this because I know that people who are engaged in hate speech and conspiracy narratives, it's not always the fact that, oh my God, there might be content moderation at some point and I need to find a coded way to communicate that I believe Jews control the White House.
It is more a game with coded language, which leads to a much more playing around with coded forms of hateful communication makes the act of communicating these ideas much more successful because you create a connection as a speaker, as a web user, as a commenter, you build a strong link between yourself and your peer group because you have some kind of a shared secretive knowledge about what is really going on. So the focus on content moderation is definitely an important angle, but I think it is even in fringe communities, you can find an incredible amount of coded language against Black people, against Jews, against Muslims, and there's no content moderation whatsoever. So it is actually something that we share as social beings that we of course want to have exclusive content even if when we demonize and dehumanize a specific target group. So these authors take this fact into account that language is permanently moving.
We're talking about the moving target. And so static lexicons therefore have a built-in temporal disadvantage and their conceptual contribution is really important. They define a task they call new hate speech terminology discovery. They distinguish it from conventional hate speech detection and from automated terminology extraction. It is not primarily a classification task, it's more kind of a discovery task. And the approach is that they provide a small set of known anchor hate terms, and then an unsupervised system searches this semantic environment around these anchors for terminology that might be associated with the same form of hate. So it's basically the so-called co-text. So it's not the overall context, historical knowledge, world knowledge, but it's actually the direct environment around these explosive terms. And importantly, it does not require a huge new manually labeled training data set every time the vocabulary changes. They test the approach across five categories, antisemitism, anti-Black hatred, anti-Muslim hatred, and anti-Asian bigotry, as well as anti-LGBTQ+ hatred.
And they evaluate the discoveries in two ways, first against established hate lexicons and second through online validation of terms that those lexicons do not contain. So this whole approach is a move from automated judgment towards augmented expertise. So the computer is very good at saying, "Here are, for instance, linguistic patterns you should investigate." And it is much less trustworthy at saying this newly discovered expression definitely constitutes hate. So a coded expression can be hateful in one community, it can be ironic in another or a reclaimed language within the targeted group itself. So the human expert, the computer actually gives you, due to the abundance of data points that we have when we talk about digital communication, a clear set of a smaller manageable data set that you as an expert can look at in order to see how a specific hateful concept or stereotype is changing over time.
And I guess this is very, very helpful. From my own experience, I can definitely say that a lot of hateful, explicitly hateful expressions show also indicate that there are more coded or novel forms of hate and exclusion in their direct environment. But of course it is just one step in the right direction because you have implicit and context depending hatred also just in the middle of non-problematic terms and phrases. And these patterns cannot be ignored, not at all, given that implicit hate, at least on the basis of my own research, sometimes occurs more than 80% within those user comments that are hateful.
Justin Hendrix:
I want to ask about one more piece in this sampler platter of different pieces that we're asking about, this little buffet that you've got on offer here in this new digital hate review. This piece that struck me in your legal category as particularly relevant at the moment, looking at AI agents and legal liability accountability, this from Sophie Xiaoyi Liu. Walk us through this, this kind of question around structural accountability absence. What's Sophie getting onto here?
Matthias J. Becker:
Yeah, so Sophie wrote a piece that actually says much of digital law assumes that somewhere behind a harmful statement, a harmful act, there is an identifiable actor to whom we can actually attach this kind of responsibility. And the problem here is that autonomous and multi-agent AI systems challenge that assumption. Liu gives that responsible actor a useful name, she calls it a legal anchor, i.e. A provider, a human decision maker towards whom obligations and liability can be directed. Her question is, what happens when harmful digital behavior no longer maps neatly onto such an actor? And she identifies two ways of losing the anchors. First is kind of the mid-operation reprogramming, an autonomous agent can acquire or execute a malicious skill after deployment. And the author draws on OWASP's 2026 Agentic Skills Top 10, including malicious skills and no governance. And the second is the autonomous swarm.
So here the author draws on recent work on malicious AI swarms, large numbers of adaptive personas capable of coordinating, infiltrating communities and manufacturing synthetic consensus with relatively little continuing human interaction. Applied to digital hate, the troubling scenario is not thousands of bots repeating the same insult. It is more the fact that there are thousands of personas that are able to adapt their language. They can coordinate harassment, they can react to moderation and make an orchestrated campaign appear quite spontaneous and that very fast. And that creates several accountability problems at once. And yeah, she lays that out in her piece beautifully and it makes it very, very transparent and accessible also for non-experts in the field.
Justin Hendrix:
I think that gives the listener a bit of a sense of what's going on in this Digital Hate Review. I hope folks will go and check it out. I'm going to make sure to include a link in the show notes, of course, and I think we'll be checking back in to see what types of new papers and new arguments and discourse come out of this particular journal in the future. Matthias, it's been great talking to you. I hope you'll come back on at some point in the future. Perhaps we can talk about what progress we've made on these issues. I'm hoping the general context changes and you will find there is more demand for the types of ideas that you are putting great effort into supply.
Matthias J. Becker:
Thank you so much for having me. Huge pleasure.
Authors


