Guild Wars Wiki 403 Forbidden

  • Recently, opening anything in Guild Wars Wiki often results in "403 Forbidden". Is this a server issue for everyone of is it somehow related to my browser (Firefox)? Usually waiting for a minute and reloading solves the issue, but it is quite annoying.

  • Nothing on my end with Firefox and Windows 11. Try using another browser. If that doesn't work, it's a general issue on your end (permissions, configuration, IP, etc, related to error 403). If it works, it's only related to your Firefox configuration.

  • From the GuildWars community on Reddit
    Explore this post and more from the GuildWars community
    www.reddit.com
    Quote

    Wiki admin here.

    Already talked to OP on discord, but posting here as well. The long story short is that the wiki (the full suite of wikis, GWW, GW2W EN, GW2W DE, GW2W FR, and GW2W ES) (and websites on the internet in general) have had to deal with a lot of botscraping over the past couple years due to the rise of LLMs, to the point that without some bordering-on-draconian firewall rules, the wikis would be too DDOS'd to be accessible to anyone.

    So if you're trying to browse special pages like WhatLinksHere and so forth, it's normal to get 403 errors while not logged in. You can solve this by making an account.

    If making an account (and being logged into your account) doesn't solve your issue, you can post on https://wiki.guildwars.com/wiki/Help:Ask_a_wiki_question or on the GW2W discord in the #wiki-not-game-support channel (just specify that it's about GW1W). Helpful information is the specific page(s) you get 403'd on, the date and time (with timezone), and your username/IP.

    On a sidenote, if you need to quickly contact GWW admins, there's a #gw1w channel on the GW2W discord that a bunch of us are in and keep an eye on.

  • Legacy also has a lot of issues with bots, we get a LOT of extra traffic lately and it is one of the reasons we had to move to a beefier server.

    But yeah, LLMs are scraping a LOT and on dynamic websites like Legacy or the GW Wiki, you do feel the impact of that enormously.

  • A bit more info was posted recently:

    TL:DR - Raise a support ticket if you have an issue / "403 error" accessing the GW Wiki.


    From the GuildWars community on Reddit
    Explore this post and more from the GuildWars community
    www.reddit.com
  • ArenaNet has been using Amazon's services for years and years for hosting everything. So when I use my personal VPN and try to access the wiki through it, I get a 403 Forbidden from it, returned from Amazon's Load Balancers.

    This most likely means that they have set up rules on their load balancers against specific providers that are often used to run AI or scraping bots. A few providers that come to mind are Hetzner, OVH, DigitalOcean and so on - not because they are necessairily bad, but because they might experience more issues with servers from those providers. The wiki itself is also mainly used by regular users and banning the IP ranges from server hosts makes sense - it'll have little to no impact for regular home internet users, but it will cut off almost all of the scraping you might experience.

    However, a lot of VPNs also utilize theze providers or the same datacenters that are being used to scrape from the wiki-servers. So it makes sense to ban them, but you will catch those on a VPN (as those have non-residential IP addresses) in the crossfire.

    There are other ways to block this kind of traffic, but on Legacy we also experience a lot of issues with scrapers as well - and the only solution I've seen that actually works against them is indeed blocking the ASN of the provider that is being used for those attacks - banning a single IP manually barely does anything for a few minutes at most.
    (An ASN-number is used to identify a provider and is linked to the IP-addresses they own. If you block an ASN number, you basically block an entire host or large chunks of their servers).
    Legacy has some special tools running on the server, like CrowdSec (and CloudFlare in front of it as well), that also try to mitigate huge scraping efforts, but those just don't cut it when an army of bots is scraping everything. They use search in rapid succession, go through posts very quickly and make a ton of requests in such a short while that it generates a huge load on our server - our server is absolutely overpowered, but such a scraping attack easily eats 20-25% of the total server resources and just block the total amount of threads that are available to PHP and the database, which ment that during such an attack, Legacy could become unavailable for regular users or just respond extremely slowly.

    AI companies say that they respect your instructions that you site should not be scraped, but that is unfortunately absolutely not the case. At this point, there's still quite a few (mostly Asian, for some reason) ASN numbers blocked on Legacy because they just generated such a load that I got constantly pinged by the monitoring services.

    (Fun fact, the main reason why Legacy started to struggle had to do with how it identifies crawlers (which search engines use to index Legacy) and it generated so many look-up queries that the database basically locked up. I took this up with the forum software developers and they have fixed this in the software.).

    I seem to get this when I don't use a VPN

    Could be that your ISP is on the ban-list, but your VPN isn't. That's a strange case though, but I don't know what rules and software they use to determine what is allowed or not.

    Hi there! I'm the Guild Wars Legacy admin, feel free to contact me if you've got issues.

    :ass: Inquisitor Karinda :der: Sunspear Elke :mes:Librarian Amber

    obey.jpg