Escaping

#1 · closed · 5 comments

View on GitHub ↗

stepankobzey

when robots.txt contains entry like this: Disallow: https://www.overstock.com/cgi-bin/d2.cgi?SEC_IID=27592&PAGE=MYACCOUNT than disallow entry is https://www.overstock.com/ and basically its skips all the urls Robots.cs line 361

Comments

sjdirect

Turns out that this entry is causing the issue... Disallow: /?PAGE=STATICPOPUP&STA_ID811 I'll try to come up with a solution soon. Thanks for reporting this.

stepankobzey

actually I got the source code and add a few fixes. split should be like this: uri.PathAndQuery.Split(new[] { '/', '?' }, StringSplitOptions.RemoveEmptyEntries); use PathAndQuery. Than there is an issue when robots has a record: Disallow: in this case its blocks entire website from crawling since its thread it as Disallow: / As well as if no robots.txt found it does not work properly. I can send you my project so you compare changes, since I download it directly

sjdirect

it would be great if you could submit a pull request. thanks!

sjdirect

Also i found this project. Not sure if its the original author or another person doing what I did. It has a nuget package to so it might be worth submitting the change to him so we could just use nuget instead of my patched version. https://code.google.com/p/robotstxt/

sjdirect

Both of these issues are now fixed an checked in with https://github.com/sjdirect/nrobots/commit/6e4134f0e253a9c12c800ec3ce1ef0b137312b46