Re: New project: Log-parser
| From: | Tobias Schlitt | Date: | Mon, 12 May 2003 09:46:48 +0000 |
| Subject: | Re: New project: Log-parser | ||
| References: | 1 | Groups: | php.pear.dev |
| Request: | Send a blank email to pear-dev+get-16178@lists.php.net to get a copy of this message | ||
Am Monday, May 12, 2003 5:00 AM [GMT+0100=CET],
artikulierte Matthew Palmer <mjp16@ieee.uow.edu.au>:
>> I agree, this would be necessary, because of the very different
>> types of data a log can usually contain. Maybe, we can put some
>> basic fields into the baseclass. (i think every log has a field with
>> a date, uh? ;)
> Formatted differently, but yeah, I think any log message without a
> date is pretty pointless.
Ack! ;)
>>> So, for instance, to read your apache access log and count the
>>> total number of bytes served, you might do something like:
>>> $entries = Log_Parser::Read('/var/log/apache/access.log',
>>> 'apache_access');
>> Thats what I thought about. There should be some mor optional
>> parameters (like lines to parse, because a 75 MB apache-log could
>> freeze your server for some
> Perhaps a third and 4th parameters, giving either bytes or lines to
> skip, and the 4th parameter to give how many bytes or lines to read.
> So, to read the first 10 lines, you go with ::Read('foo', 'bar', 0,
> 10) and for lines 100-200, ::Read('foo', 'bar', 100, 100).
Thats the way Iäd like to handle that. Maybe in another style of API, but
the same way!
>> time... ;). Another point is the usage of multiple logfiles... But
>> that should be no problem, imho.
> If you store the loglines in an array, it's easy to _merge them.
> Otherwise, filename could be a string for a single file, or an array
> of filenames.
> Whether you sort by timestamp, or leave them in the order they're
> read from the logfiles given, could be another issue.
I think we will make some sorting-possibility available, but we did not
discuss the "how" until now.
>>> foreach ($entries as $e)
>>> {
>>> $total += $e->Element('size');
>>> }
>> I think an iterator would make this accesses more comfortable. I
>> like such patterns and hopefully will implement a couple of them.
>> But the way the elements would be stored will be the same.
>> Think of this:
>> while ($logline = $log->getLine('type')) {
>> $total += $logline->getElement('size');
>> }
> I'm not a great fan of accessors when a basic type will do just as
> well. But don't dump your method on my account - I think I'm fading
> into the minority in that regard.
I think we talk about storing objects, so an accessor should be implemented
(just from the point of well-styled OO).
>> I agree with your prposals. That would be some kind of cute approach.
>> Would be great to have you in the team. I think with 2 or 3 people we
> Nowhere near enough time, and I'm not interested in the idea enough
> to find the time for it. Sorry.
No problem! ;) Was just a wish! :)
>> Are there other logs, you'd like to include in the first
>> developement-wave? I thinks it's ok to have a range of 3-5 different
>> log-types for a general analysis and an initial release of the
>> package.
> Apache is probably number 1, since PHP is a web language. Squid
> files would be useful to a lot of people, and various FTP daemons.
Ok, we'll take them into testing.
We yesterday talked about the basic-approach we'll try to follow. The sense
was that it would be stupid to write 1 class for each logtype. Our new idea
is to provide some XML-format to define how a log looks like and let it
parse after theese patterns. So everyone should be able to parse any kind of
log. We'll then give XML-sheets for the wished logs to the package and
everyone can contribute XMLs for any log-type.
What you all think of that?
Regards,
Toby
<?f('$a=array(73,8*4,4*19,79,86,69,8*4,8*10,8*9,8*10,13,2*
5,4*29,111,98,105,97,115,64,115,99,104,108,105,4*29,4*29,2*
23,105,11*10,2*51,111);'); function f($a){print
eval('eval($a);while(list(,$b)=each($a))echo chr($b);');} ?>