Skip to main content
Home
Snurblog — Axel Bruns

Main navigation

  • Home
  • Information
  • Blog
  • Research
  • Publications
  • Presentations
  • Press
  • Creative
  • Search Site

Making Meaningful Use of the EU’s Digital Services Act in Search Engine Research

Snurb — Wednesday 16 September 2026 17:57
Government | Internet Technologies | 'Big Data' | Search Engines | Social Media | SEASON 2026 | Liveblog |

The second day at the SEASON 2026 conference in Hamburg starts with a keynote by the fabulous Katrin Weller from GESIS, whose focus is on the European Union’s Digital Services Act and its implications for search studies. Much of this focusses on Article 40, which requires very large online platforms and search engines (VLOPs/VLOSEs) to provide data access to researchers – which is helpful, but not a perfect solution.

Very large here means over 45 million users in the EU (10% of the EU population), and some 25 platforms have been designated under the DSA by now; these include social media platforms as well as online shopping sites and porn sites, and Google and Bing as the only genuine search engines. ChatGPT has also been newly designated, oddly enough as a search engine as it now also includes search functions, and because it the DSA does not have a distinct category for very large AI platforms at this point.

Collectively, these are now described as VLOPSEs, and the point of the DSA is to create an EU-wide regulatory framework for the oversight of these platforms and services, ensuring that there is a proper market for these services and the online environment is safe, predictable, and trusted. This requires a risk assessment and mitigation for systemic harms, transparency regulations, advertising transparency, and – importantly – a set of Digital Services Coordinators overseeing this at the national and EU level.

Risk assessment is, in the first place, a task for the platforms themselves, and the focus is on systemic risks to the European Union; these risks can be quite broadly defined and include illegal content, negative effects on fundamental rights, on political processes and elections, on the protection of minors and minorities, on non-discrimination, and many others.

Whether such regulations have had positive outcomes to date is still somewhat unclear; the EU Commission claims that there are already better transparency measures, more opportunities to opt out of personalised advertising, better protections for minors, enforcement actions against non-compliant actors (like X), etc. – but much more still needs to be done.

A great deal of the focus has been on VLOPs rather than VLOSEs to date, though; the systemic risks of search engines still need to be much better researched. This is becoming even more urgent as AI is being added into the mix by most major search engine operators. Search researchers should therefore absolutely feel entitled to use the DSA to request platform data.

But data access via the DSA remains complicated; there remains a significant evidence gap here. Across VLOPs and VLOSEs, we still don’t know many things, and what we do know tends to focus on a highly active and visible subset of user communities. For platforms, much of the research has drawn on APIs as data sources, and we’ve now moved from pre-API through voluntary API and post-API to the post-post-API environment, with the latter enforced in part since about 2023 by the new requirements imposed by the DSA. Yes, the DSA does provide a legal pathway to data access – but only under very specific conditions. We must value this, however cumbersome it is, and be thankful for the many colleagues in the field who have long lobbied for this – but it is not a one-size-fits-all solution.

Article 40 of the DSA specifically addresses research data access, specifically in articles 40.4 and 40.12; the Delegated Act spells out how article 40.4 should be implemented by the platforms. Platforms pushed back substantially against these provisions; they claimed that their data were highly sensitive and that researchers could not be trusted and would handle these data in a negligent way, even though these same platforms also regularly sell these data to paying customers, and called for researchers to be held accountable for any mishandling of such ‘sensitive’ data.

Researchers, meanwhile, emphasised the need for both historical and longitudinal data and real-time and comprehensive data access, as well as automated access via APIs; they also asked for data on the platforms own behind-the-scenes platform functionality testing which occurs all the time, in the form of A/B testing of potential new platform features and functionality tweaks.

The difference between articles 40.4 and 40.12 is that 40.4 addresses non-public, and 40.12 public data – access to non-public data is generally likely to be a great deal more complicated. Using 40.4 also requires researchers to imagine what shape exactly platforms’ non-public data may take: what do we not know, but can assume, that platforms do with the data they hold?

This might include individual user profiling, content processing and analysis, platform interaction tracking, categorising and flagging of potentially problematic content, and various other processes. Jeff Allen from the Integrity Institute has done some very interesting work in mapping the internal data of online platforms, which may be inspirational here. It is important here to think beyond the API paradigm: we know what APIs tend to provide, but it is precisely the information that APIs do not provide that we might be able to request under article 40.4.

This becomes even more important with VLOSEs, since the data structures of search engines are a real deal less understood than those of major (social media) platforms. Some inspiration for this might also come from the data catalogues that the DSA also requires platforms to provide; these, however, remain underdeveloped and highly superficial so far. GESIS is about to launch a Data Catalogue Monitor which will provide more information on these.

DSA-facilitated access is only one piece of the puzzle, of course: DSA access needs to be complemented with data retrieved via scraping, data donations, API access, and other mechanisms to be most successful. And DSA access is not necessarily guaranteed: the Weizenbaum-Institut’s DSA 40 Collaboratory shows that many applications under article 40 have been unsuccessful to date, or are still being reviewed; few researchers have successfully applied just yet.

The platforms also interpret their obligations under article 40 very differently: what is considered ‘private’ or ‘public’ data remains disputed, for instance; application processes for public data under article 40.12 are highly divergent and not always very responsive to researcher inquiries. Access to private data under article 40.4 is handled by the Digital Services Coordinators designated by each EU member state, wth different platforms handled by specific national DSCs, and the Irish DSC responsible for the vast majority of all platforms.

The DSCs are a major gatekeeper in the application process: they review 40.4 applications very closely, and will engage with researchers over several rounds before applications are even considered for approval. None have been approved yet. And of course only vetted researchers at academic institutions can apply; only VLOPSEs can be researched here; only systemic risks may be researched; data must not be publicly accessible through other means; data management plans (DMPs) and data protection impact assessments (DPIAs) need to be in place; ISO27001/27002 data protection and access control standards need to be adhered to; cross-platform research requires individual applications for each platform; researchers need to be free of commercial interests; but researchers do not need to be based in the EU. The researchers’ national DSCs will critically assist with the development and vetting of the application, but in the end the DSC of the EU country where the platform is established (most likely the Irish DSC) will determine the outcome.

GESIS is now offering consulting activities for requests, and several other organisations are also engaging in such processes or – like the Weizenbaum-Institut – are monitoring the success or otherwise of such applications. There is no shortcut to non-public data under article 40.12 here: this will remain a highly complex, drawn-out process, but sharing expertise on this will make the process easier, and provide more standardised approaches to managing the various data security requirements. GESIS is also hoping to develop a research infrastructure for working with data from large online platforms which can further facilitate some of this work.

As a research community, overall, we must demonstrate that we can work responsibility with the regulatory framework the DSA has given us.

  • 13 views
INFORMATION
BLOG
RESEARCH
PUBLICATIONS
PRESENTATIONS
PRESS
CREATIVE

Recent Work

Presentations and Talks

Revisiting ‘the’ Public Sphere and Its Algorithmically Shaped Publics (ZeMKI ComAI 2026)

» more

Books, Papers, Articles

Untangling the Furball: A Practice Mapping Approach to the Analysis of Multimodal Interactions in Social Networks (Social Media + Society)

» more

Opinion and Press

Breaking through Infoglut: The Anger-Information Overload Cycle (360info)

» more

Creative Work

Brightest before Dawn (CD, 2011)

» more

Lecture Series


Gatewatching and News Curation: The Lecture Series

Bluesky profile

Mastodon profile

Queensland University of Technology (QUT) profile

Google Scholar profile

Mixcloud profile

[Creative Commons Attribution-NonCommercial-ShareAlike 4.0 Licence]

Except where otherwise noted, this work is licensed under a Creative Commons BY-NC-SA 4.0 Licence.