Re: parse HTML page

From: Date: Tue, 25 Jul 2000 15:45:12 +0000
Subject: Re: parse HTML page
References: 1  Groups: php.general 
Request: Send a blank email to php-general+get-8065@lists.php.net to get a copy of this message
""Erb, Maria"" <merb@keene.edu> wrote in message news:CFD3579B7DE2D111A7C000805F9F61AD0162A961@dinsmore.keene.edu... > hi, > > I need to be able to parse an HTML page and extract the URL's from it. I am > really unsure how to go about doing this. Any suggestions? > I had to do this recently. fgetss() is useful for this, as are regular expressions (but I can't figure regexps out for the life of me). Try this set of steps ... - Open the file for input, and perhaps another file for output (or use an array for output) - Strip all of the extraneous tags from each line using $string = fgetss($file,4096,"<a></a>") - In the same loop, check each string for "<a"; if one is found, search for the matching "</a>" - Use the appropriate string functions to get this link into a single string (functions like substr, strpos, etc), then put that string in your array. - You can write additional functions with regexp's or plain ol' string routines to get info like the link itself, the description, and other information included in a <A> tag. Note that a single regexp could probably do the work of the 2nd, 3rd, and 4th steps. -- O.

« previous php.general (#8065) next »