Has Zenodo become the place where "citizen scientists" post their Ai generated papers since Arxiv put the vouching in place? I've seen more HN post from this site to questionable papers
There is no moderation, so, yes, anything (legal) goes, I suppose. But to the point: it is interesting that even CERN is taking a hit from the LLM crawlers.
I suspect that the growth in LLM "crawlers" is part traditional crawlers and a growing part WebFetch tools agents use directly on behalf of users.
Personally, my page fetch count is way up because I have a custom deep research agent, I've almost weaned myself off traditional search, but that agent is fetching dozens of pages, like I used to click through a bunch of search results to figure out which are relevant, adjust my query, etc...
That now happens at a higher frequency and volume through my agent.
More people can serve up static markdown for the agents with `Content-Type: application/markdown`, CDNs work really well. I think both sides of this dynamic would find it beneficial.
I suspect much of the infra stress is on all the "enhancements" found in the modern web application and the targeted advertising ecosystem
Has Zenodo become the place where "citizen scientists" post their Ai generated papers since Arxiv put the vouching in place? I've seen more HN post from this site to questionable papers
There is no moderation, so, yes, anything (legal) goes, I suppose. But to the point: it is interesting that even CERN is taking a hit from the LLM crawlers.
I suspect that the growth in LLM "crawlers" is part traditional crawlers and a growing part WebFetch tools agents use directly on behalf of users.
Personally, my page fetch count is way up because I have a custom deep research agent, I've almost weaned myself off traditional search, but that agent is fetching dozens of pages, like I used to click through a bunch of search results to figure out which are relevant, adjust my query, etc...
That now happens at a higher frequency and volume through my agent.
> agents use directly on behalf of users
Could well be, who knows? But, well, then, in a sense you're a part of the problem, causing a DoS for infrastructures and other, human users.
More people can serve up static markdown for the agents with `Content-Type: application/markdown`, CDNs work really well. I think both sides of this dynamic would find it beneficial.
I suspect much of the infra stress is on all the "enhancements" found in the modern web application and the targeted advertising ecosystem