What people do with their own domains is their own business, but
deb.devuan.org is the Devuan teams business, it's our domain. The Devuan
team decides what goes on
http://deb.devuan.org/ The same applies to the
CC.deb.devuan.org domains.
So all Devuan package mirrors should be serving the exact same thing when
responding to
http://deb.devuan.org/ and the country code versions.
We have
https://pkginfo.devuan.org/ for searching for packages, so we
don't really need the search bots scanning us. The AI bots found non
Devuan things on some of our mirrors at
http://deb.devuan.org/, hence why
they keep scanning
http://deb.devuan.org/ for non Devuan package repos.
So we decided to just stop the bots from scanning
http://deb.devuan.org/
which is why there is now a
http://deb.devuan.org/robots.txt that tells
all the bots that bother to obey it not to scan
http://deb.devuan.org/
For those bots that ignore robots.txt there will soon be a
http://deb.devuan.org/index.html file. This will include links to the
package repo stuff and no other links. "
http://deb.devuan.org/" should
serve this index.html.
In this way the behaving bots just wont scan anything, and the
misbehaving bots will only find Devuan package repo, nothing else, when
they scan
http://deb.devuan.org/
Hopefully within a week the bots will have forgotten the non Devuan stuff
they have been scanning
http://deb.devuan.org/ for, and they wont be
bogging down our mirrors and filling our logs with 404s. I'm certainly
tired of the same non existent files being scanned for 7000 times a day.
This should also help with the bots scanning Devuan package mirrors for
non existant packages, that are not at deb.devuan.org, but that can still
leak a bit.
Cross our fingers, hope this helps.