RE: [PHP] ereg / regexp fun

From: Date: Wed, 16 Aug 2000 18:02:37 +0000
Subject: RE: [PHP] ereg / regexp fun
Groups: php.general 
Request: Send a blank email to php-general+get-12080@lists.php.net to get a copy of this message
Hi Regex Experten, Maybe there is someone who can solve my problem as well ? ;-) For a scientific project which involves medical vocabularies I need to harvest a large number of Websites for medical terms. I have managed to program a kind of meta search engine which retrieves all links offered by Google.com for a given search term. It puts the links into a database for further processing. My next step before actually "reading" the files is that I want to get all child links from the stored URLs. So I am trying to extract all links from these pages. My function should return all links in the given URL as an array. However as you will see my lack of regex knowledge prevents it from working. Any helping idea? Heiko All URLs come from the database /* test URL */ $url = "http://dentistry.vh.org/sites.html"; function harvestlinks($url){ $i = 1; /* Opening and reading the file */ $file = fopen($url, "r"); $rf = fread($file, 150000); /* Extracting all matches of a link regardless of local or remote, need later to take care of local links */ if(preg_match_all("<a href=\"http://[[:alpha:]]\"", $rf, $matches)){ while (list($key, $val) = each($matches)) { print "$i.link=>$val<BR>"; // just print them for checking $i++; } } } Or, maybe someone knows of an available snippet extracting all links and convert them to ones with absolute path??? Thanks! Heiko -- Heiko Spallek, DMD, Ph.D.: heiko@spallek.com Asst. Professor Department of Dental Informatics Temple University School of Dentistry Try: http://www.temple.edu/dentistry/di/

« previous php.general (#12080) next »