SFU.CA

Managing AI Traffic for PKP Hosted Journals

Feature image for post. Text reads the same as the title of the post. The PKP logo is included. The photograph used for the background is a Sonora Desert prickly pear cactus in bloom with a pink flower surrounded by inch long spikes, representing the protection of PKP hosted journals from corporate AI interference.
“Prickly Pear Cactus in Bloom” photo in Sonora Desert by PKP’s Famira Racy. Flowers represent PKP hosted journals using OJS and spikes represent PKP working to protect those journals from commercial AI traffic interference.

PKP’s Head of Systems, Michael Felczak, underlines how commercial AI traffic has been putting a strain on hosting providers such as PKP. Read on to learn how PKP is balancing traditional and new tools to protect its hosted journals.

With the rapid growth of commercial AI services and a plethora of applications and agents built upon these services, hosting providers have witnessed an exponential growth in network traffic to their hosted websites. Automated traffic has recently surpassed human-generated traffic and today represents the bulk of access requests to online journals, monograph collections, and preprint servers. 

Keen to train their LLMs, AI companies and service providers view open access scholarly publishing as a free source of data that does not require permission or contractual arrangements typical among commercial operators. 

As a result, hosting providers are struggling to maintain an acceptable level of service for readers due to new demands from automated traffic, which places high demands on server and network resources. Much of this automated traffic is aggressive, involving requests for pages and URLs at high rates that far exceed human access to published content.

Protecting PKP Hosted Journals

To address the problem of automated traffic, PKP Publishing Services has deployed traditional network management tools alongside new tools that target AI harvesting. These traditional methods include firewalls to restrict origin IP addresses, network throttling to slow down aggressive traffic, and community resources that aggregate information about offending service providers. 

While traditional methods are effective to some extent, they are limited to a set of rules that must establish in advance what and who to block. Today AI crawlers often adjust their behaviour, conceal their fingerprints, and switch to a different set of origin IP addresses. 

The end result is an AI-based iteration of whack-a-mole that requires staff time and resources to monitor traffic, adjust rules, and deploy additional filters to maintain an acceptable level of service for hosted journals and readers. Equally importantly, since AI services are now commonplace for discovery of content and search results, commercial AI services require access to published content and cannot be completely blocked. This additional AI-centric traffic and server activity requires additional server and networking resources.

Looking Ahead

Traditional network management tools co-exist alongside new tools and services that target AI automated traffic. Recently several open source projects have been launched that support dynamic AI crawler filtering and PKP Publishing Services has started to deploy these tools on a limited scale to assess their effectiveness. 

Anubis, the most popular of these projects, filters human traffic from automated traffic by presenting a challenge to website visitors, similar to Google’s reCaptcha. Unlike Google’s reCaptcha, the Anubis challenge does not require user input and is instead verified by the user’s browser. 

Even if AI bots adjust their behaviour or origin IP addresses, if they are unable to complete the Anubis challenge they will continue to be actively blocked. Since it doesn’t require user input, from an accessibility perspective the Anubis challenge is more accessible compared to a traditional captcha. 

This ongoing work represents the commitment of PKP Publishing Services to protect hosted journals from aggressive traffic while balancing the need of emerging AI services to index open access content to ensure that it can be discovered by readers.