Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Search results are also not supposed to be indexed. I believe this is actually in the guidelines.


Yup. Search for [quality guidelines] and they're at http://www.google.com/support/webmasters/bin/answer.py?answe... . The relevant part is "Use robots.txt to prevent crawling of search results pages or other auto-generated pages that don't add much value for users coming from search engines."


Matt, I tried that recently with one of my own sites, and after a few days I got a warning in my Google Webmaster Tools dashboard that the bot could not access my /search URL because it was blocked in my robots.txt, and I should take action to correct it. So I then unlocked my /search URL, which is in violation of the above guideline, but made the error in Google Webmaster Tools go away.

These conflicting messages from Google are very confusing!


Hi j_col-

In that case, Google Webmaster tools is not actually reporting an error. That's a report to show you what URLs Google tried to crawl but couldn't (due to being blocked) so you can review it and ensure that you are not accidentally blocking URLs that you want to have indexed.

I agree that it's confusing in that the report is in the "crawl errors" section.

(I built Google webmaster tools so this confusion is entirely my fault; but I don't work at Google anymore so sadly I can't fix this.)


Thanks for the response, I will block my /search URL once more via robots.txt and will ignore the warnings in the Webmaster Tools.


Literally Google Webmaster Guidelines say: "Use robots.txt to prevent crawling of search results pages or other auto-generated pages that don't add much value for users coming from search engines."

So a couple of questions: - What is of value for a user? - Who is determining the value for the users in these cases?

As always, it's not always quite clear on what treatment you should use for search pages!


I'm pretty sure that having different search engines indexing each other would be bad. Don't cross the streams.


If I recall correctly, crossing the streams did kill marshmallow man!

In any case, I agree with you, search engine indexing search results would be bad, but the line is not that clear all the time!

Some vertical search engine result pages are a great and relevant result from a user perspective on the question they are trying to solve.


It's hard to draw the line on what is search and what isn't. If you use a system where the last part of the URL will be treated as a search but only certain ones are ever linked to so they are a product page.

Also if you have some kind of recent searches list and those link.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: