Re: parse HTML page
| From: | O. | Date: | Tue, 25 Jul 2000 15:45:12 +0000 |
| Subject: | Re: parse HTML page | ||
| References: | 1 | Groups: | php.general |
| Request: | Send a blank email to php-general+get-8065@lists.php.net to get a copy of this message | ||
""Erb, Maria"" <merb@keene.edu> wrote in message
news:CFD3579B7DE2D111A7C000805F9F61AD0162A961@dinsmore.keene.edu...
> hi,
>
> I need to be able to parse an HTML page and extract the URL's from it. I
am
> really unsure how to go about doing this. Any suggestions?
>
I had to do this recently. fgetss() is useful for this, as are regular
expressions (but I can't figure regexps out for the life of me). Try this
set of steps ...
- Open the file for input, and perhaps another file for output (or use an
array for output)
- Strip all of the extraneous tags from each line using $string =
fgetss($file,4096,"<a></a>")
- In the same loop, check each string for "<a"; if one is found, search for
the matching "</a>"
- Use the appropriate string functions to get this link into a single string
(functions like substr, strpos, etc), then put that string in your array.
- You can write additional functions with regexp's or plain ol' string
routines to get info like the link itself, the description, and other
information included in a <A> tag.
Note that a single regexp could probably do the work of the 2nd, 3rd, and
4th steps.
--
O.