Re: Search engine recommendation (htdig vs. udmsearch vs.glimpse)???
| From: | Rasmus Lerdorf | Date: | Thu, 27 Jul 2000 15:55:45 +0000 |
| Subject: | Re: Search engine recommendation (htdig vs. udmsearch vs.glimpse)??? | ||
| References: | 1 | Groups: | php.general |
| Request: | Send a blank email to php-general+get-8572@lists.php.net to get a copy of this message | ||
> > I have used all three over the years. My current favourite is
> > udmsearch. There will probably be some native support for udmsearch
> > showing up in PHP soon.
>
> I looked at it and it seems to use a database (as in MySQL) as a
> backend. Does this really scale well? I mean, is it possible to do a
> phrase or boolean search on a 100MB database in the same speed as, say,
> htdig or swish/swish++?
It is actually really fast. When run in crc-multi mode, it only stores
the crc32 checksum of the words in a bunch of small tables. The
distribution across tables is based on word length. So when doing a
search you get indexed lookups in fixed-size tables. In MySQL at least
this is really really fast.
For me what killed htdig was its indexing speed. I am indexing a very
busy mail archive where messages flow in quickly. On a Dual PII-550 box
htdig actually indexed slower than new data was flowing in. That made it
impossible for me to use htdig. UdmSearch is an order of magnitude faster
at indexing. I am not sure why htdig is so slow. It starts out pretty
fast, but once your indexes grow large it slows down significantly.
Another cool thing about UdmSearch is that you can make intelligent
queries. For example, if I want to know how many urls have been indexed,
and how many total bytes, I can simply do:
mysql> select count(url), sum(docsize) from url;
+------------+--------------+
| count(url) | sum(docsize) |
+------------+--------------+
| 172848 | 1065051776 |
+------------+--------------+
1 row in set (4.70 sec)
So I currently have 172,848 documents indexed for a total of about a Gig
worth of data. Searches, even complex ones, on this 1 Gig database are
quick. For example, I just tried a search of:
rasmus & ~lerdorf
That is, find all documents that have my firstname in them, but not my
lastname. It returned 2451 documents in under a second.
And of course, repeated queries for the same thing will be faster because
the database backend will cache query results.
As for the PHP portion of this message... ;) I added a crc32() native
function to PHP 4.0.1 to eliminate a system call from the search
interface. I plan on adding support for libudmsearch to speed up the
PHP-driven search interface even further.
-Rasmus