Re: extract data from html
| From: | Hugh Bothwell | Date: | Sat, 30 Jun 2001 01:26:17 +0000 |
| Subject: | Re: extract data from html | ||
| References: | 1 2 | Groups: | php.general |
| Request: | Send a blank email to php-general+get-55617@lists.php.net to get a copy of this message | ||
> > > 1. Open the html file in read only mode
> > > 2. Start reading the html file till I encounter a <td> tag (I don't
know
> > > how to do this)
> > > 3. Grab that data after the <td> tag (and then what?)
> >
> > See <http://php.net/manual/en/function.fopen.php>
> > and
> > <http://php.net/manual/en/function.fgetss.php> plus
> > the chapter for
> > whatever DBMS you want to drop the file contents into.
>
> Thanks. One thing just reading the manual without the idea of how the
> function works is of no use. Some examples would help. In fact I did use
> fopen, fgets, fgetss but the problems is that the html tag that I am
> looking is <td>. Now this is easy but if <td width="25%"> or <td
> colspan="7"> would give a problem.
(grimace) The PHP manual is actually very well written; I can usually find
exactly what I need in < 10s. I think your complaint just covers sloppy
thinking.
I'd think you should be able to find screen-scraper code around; if not, try
this:
- search for '<td'. NOTE: use a case-insensitive search!
- search for the first trailing '>'. Save (this character position + 1).
- search for the first trailing '</td>'. Again, case-insensitive!
- store everything between the two; strip all HTML tags, add slashes, and
store it.
- increment your file position by 5 characters and repeat.
I'd give you actual code, but I think you could use some manual practice
(smirk).