Re: Y! BOSS Search
| From: | Hannes Magnusson | Date: | Sat, 02 Jan 2010 21:41:09 +0000 |
| Subject: | Re: Y! BOSS Search | ||
| References: | 1 | Groups: | php.webmaster |
| Request: | Send a blank email to php-webmaster+get-6890@lists.php.net to get a copy of this message | ||
On Sat, Jan 2, 2010 at 20:45, Stewart Lord <stewey@ambitious.ca> wrote:
> Hi Hannes & Philip,
>
> I have been researching Yahoo Boss, and it looks like there are a number of
> ways we could improve on the existing search (I'm not referring to the
> autocomplete search, but rather the full-page search).
>
> I found an interesting article on Tech Crunch about how they implemented
> Yahoo Boss:
>
> http://www.techcrunch.com/2008/11/26/techcrunchs-new-search-engine-powered-by-yahoo-boss/
>
> One of the key things that they do is manually feed data to Yahoo via XML.
> This allows them to associate arbitrary metadata with each page. That way
> they can offer advance search options like search by author or category.
>
> Are we feeding any additional data to Yahoo, or just rely on them to crawl
> the sites? If it's the latter, would we want to start manually feeding data
> doing to provide better search options? How would that work with our
> infrastructure?
All events, conferences, news, persons and manual pages are markedup
with semantical markup using bunch of different microformats and eRDF
markup. Furthermore people.php.net uses RDFa for the person profiles,
while individuals in the manual and news.php.net use hcard.
There are currently 3 different SearchMonkey applications to extract
information about the manual pages[1].
We could write more applications to present the news, events and
conferences information in more fun way, and could potentially do
something interesting to the changelogs and download pages.
I don't think we should feed Y! anything different then any other
search engine (and we don't currently, with the exception of
news.php.net, where we block all search engines other then Y!). We can
present the information in fun ways, but if the search engine isn't
good enough to extract our information we should rather pick another
search engine.
Providing additional data via in pure RDF or DataRSS however is an
option, thats something everyone can use.
-Hannes
[1] http://gallery.search.yahoo.com/search?p=php