Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I don't completely disagree in that perhaps Bing should respect Robots.txt when using information gathered in this way. However, the aim of Robots.txt is to restrict the actions of a automated web-crawler (bandwidth limiting and preventing unwanted interaction with dynamic pages), in this case the crawling is not carried out by a 'robot' so in my opinion there is room for interpretation.

I guess its an interesting point, clearly denying access to an area in the robots.txt suggests you do not have permission to use that information. However, only the TOS will be definitive on what you can and cannot do. For example not including a robots.txt file clearly does not waive all rights, but interpreting the TOS is clearly beyond the capability of an automated system.

In this case I would suggest the only safe course of action for Bing would be to have an exclude list of domains and allow anyone to have their own site excluded from this information gathering.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: