Re: Search engine recommendation (htdig vs. udmsearch vs.glimpse)???

From: Date: Thu, 27 Jul 2000 15:55:45 +0000
Subject: Re: Search engine recommendation (htdig vs. udmsearch vs.glimpse)???
References: 1  Groups: php.general 
Request: Send a blank email to php-general+get-8572@lists.php.net to get a copy of this message
> > I have used all three over the years. My current favourite is > > udmsearch. There will probably be some native support for udmsearch > > showing up in PHP soon. > > I looked at it and it seems to use a database (as in MySQL) as a > backend. Does this really scale well? I mean, is it possible to do a > phrase or boolean search on a 100MB database in the same speed as, say, > htdig or swish/swish++? It is actually really fast. When run in crc-multi mode, it only stores the crc32 checksum of the words in a bunch of small tables. The distribution across tables is based on word length. So when doing a search you get indexed lookups in fixed-size tables. In MySQL at least this is really really fast. For me what killed htdig was its indexing speed. I am indexing a very busy mail archive where messages flow in quickly. On a Dual PII-550 box htdig actually indexed slower than new data was flowing in. That made it impossible for me to use htdig. UdmSearch is an order of magnitude faster at indexing. I am not sure why htdig is so slow. It starts out pretty fast, but once your indexes grow large it slows down significantly. Another cool thing about UdmSearch is that you can make intelligent queries. For example, if I want to know how many urls have been indexed, and how many total bytes, I can simply do: mysql> select count(url), sum(docsize) from url; +------------+--------------+ | count(url) | sum(docsize) | +------------+--------------+ | 172848 | 1065051776 | +------------+--------------+ 1 row in set (4.70 sec) So I currently have 172,848 documents indexed for a total of about a Gig worth of data. Searches, even complex ones, on this 1 Gig database are quick. For example, I just tried a search of: rasmus & ~lerdorf That is, find all documents that have my firstname in them, but not my lastname. It returned 2451 documents in under a second. And of course, repeated queries for the same thing will be faster because the database backend will cache query results. As for the PHP portion of this message... ;) I added a crc32() native function to PHP 4.0.1 to eliminate a system call from the search interface. I plan on adding support for libudmsearch to speed up the PHP-driven search interface even further. -Rasmus

« previous php.general (#8572) next »