Re: Y! BOSS Search

From: Date: Sat, 02 Jan 2010 21:41:09 +0000
Subject: Re: Y! BOSS Search
References: 1  Groups: php.webmaster 
Request: Send a blank email to php-webmaster+get-6890@lists.php.net to get a copy of this message
On Sat, Jan 2, 2010 at 20:45, Stewart Lord <stewey@ambitious.ca> wrote: > Hi Hannes & Philip, > > I have been researching Yahoo Boss, and it looks like there are a number of > ways we could improve on the existing search (I'm not referring to the > autocomplete search, but rather the full-page search). > > I found an interesting article on Tech Crunch about how they implemented > Yahoo Boss: > > http://www.techcrunch.com/2008/11/26/techcrunchs-new-search-engine-powered-by-yahoo-boss/ > > One of the key things that they do is manually feed data to Yahoo via XML. > This allows them to associate arbitrary metadata with each page. That way > they can offer advance search options like search by author or category. > > Are we feeding any additional data to Yahoo, or just rely on them to crawl > the sites? If it's the latter, would we want to start manually feeding data > doing to provide better search options? How would that work with our > infrastructure? All events, conferences, news, persons and manual pages are markedup with semantical markup using bunch of different microformats and eRDF markup. Furthermore people.php.net uses RDFa for the person profiles, while individuals in the manual and news.php.net use hcard. There are currently 3 different SearchMonkey applications to extract information about the manual pages[1]. We could write more applications to present the news, events and conferences information in more fun way, and could potentially do something interesting to the changelogs and download pages. I don't think we should feed Y! anything different then any other search engine (and we don't currently, with the exception of news.php.net, where we block all search engines other then Y!). We can present the information in fun ways, but if the search engine isn't good enough to extract our information we should rather pick another search engine. Providing additional data via in pure RDF or DataRSS however is an option, thats something everyone can use. -Hannes [1] http://gallery.search.yahoo.com/search?p=php

« previous php.webmaster (#6890) next »