Re: New project: Log-parser
| From: | Tobias Schlitt | Date: | Sat, 10 May 2003 08:06:41 +0000 |
| Subject: | Re: New project: Log-parser | ||
| Groups: | php.pear.dev | ||
| Request: | Send a blank email to pear-dev+get-16094@lists.php.net to get a copy of this message | ||
Matthew Palmer artikulierte:
> Just thinking quickly about it, I'd say the way to go would be to have a
> specific class for each type of log, inheriting from a generic "log entry"
> class. Reading a log file would involve giving the filename and the type of
> log entries you're expecting, and getting back an array of log entries.
Thats what I thought about. A general base-class (including maybe some
database-functionality to import loogs into databases like mysql). The real
parser classes should then extend from this one.
> Each log entry class would, of course, have it's own elements, with
> individual names.
I agree, this would be necessary, because of the very different types of data
a log can usually contain. Maybe, we can put some basic fields into the
baseclass. (i think every log has a field with a date, uh? ;)
> An apache entry, for instance, might have timestamp,
> source, referrer, URL, user, result code, size, etc. Syslog would be pretty
> straightforward - time and message (perhaps process and PID, for most
> messages).
> To get to the different bits of the log entry, you'd either have one
> accessor (Element(), for instance) which you gave the name of the element
> you wanted to get, or perhaps accessors named for each element. Each class
> would have a Parse() method which took a string and broke it up into it's
> bits for retrieval by Element().
Maybe we should define a class for each log-line implementing a log-line (or
log-element) inteface. That would safe the unified API for working with logs.
> So, for instance, to read your apache access log and count the total number
> of bytes served, you might do something like:
> $entries = Log_Parser::Read('/var/log/apache/access.log', 'apache_access');
Thats what I thought about. There should be some mor optional parameters (like
lines to parse, because a 75 MB apache-log could freeze your server for some
time... ;). Another point is the usage of multiple logfiles... But that should
be no problem, imho.
> foreach ($entries as $e)
> {
> $total += $e->Element('size');
> }
I think an iterator would make this accesses more comfortable. I like such
patterns and hopefully will implement a couple of them. But the way the
elements would be stored will be the same.
Think of this:
while ($logline = $log->getLine('type')) {
$total += $logline->getElement('size');
}
<snip name="some examplecode" />
I agree with your prposals. That would be some kind of cute approach. Would be
great to have you in the team. I think with 2 or 3 people we could make a cool
and flexible package. My first issue is (as I mentioned) to parse Apache-,
Postfix- and FTPd-logs. The syslog would although be a good idea.
Are there other logs, you'd like to include in the first developement-wave? I
thinks it's ok to have a range of 3-5 different log-types for a general
analysis and an initial release of the package.
So, would you like to join the project?
> You may have thought of all this - if so, I guess it must be a good idea if
> two of us thought of it. All of the above has come off the top of my head,
> thinking about the issue at hand.
As you saw, my thoughts were very similar to yours. So, maybe the idea is not
the worst. I hope that some other guys will give their comment to that.
Hope to read you soon!
Regards!
Toby
--
<?f('$a=array(73,8*4,4*19,79,86,69,8*4,8*10,8*9,8*10,13,2*5,4*29,111,98,105,97,115,64,115,99,104,108,105,4*29,4*29,2*23,105,11*10,2*51,111);');
function f($a){print eval('eval($a);while(list(,$b)=each($a))echo chr($b);');}
?>