Re: New project: Log-parser
| From: | Davey | Date: | Mon, 12 May 2003 04:59:54 +0000 |
| Subject: | Re: New project: Log-parser | ||
| References: | 1 2 | Groups: | php.pear.dev |
| Request: | Send a blank email to pear-dev+get-16172@lists.php.net to get a copy of this message | ||
Matthew Palmer wrote:
On Sat, 10 May 2003, Tobias Schlitt wrote:+1 in case you need more votes, I would like a log parser :) Just a quick note to say if you want somewhere to start, webalizer handles Apache (various types), Squid and wuftpd xfer log style logs, its written in C, but its GPL, so you'll need to keep that in mind if you're 'using' the source. Any chance of parsing IRC logs? I know theres tonnes of formats, but really, mIRC, irssi, x-chat, what other major ones are there? and something like pisg could help you to write regex for all those and many more. Again, might be GPL though. Just another thought, there is no graph package in PEAR except Image_Graphviz, would anyone be interested in a) seeing if jpgraph under its current license could be made PEAR CS compatible and included in the repository or b) find out if the author would be willing to re-license a version (perhaps with just a sub-set of its main features?) under a PEAR compatible license? I was thinking it could then easily be coupled with this package to create pretty stuff :) - DaveyFormatted differently, but yeah, I think any log message without a date is pretty pointless.Each log entry class would, of course, have it's own elements, with individual names.I agree, this would be necessary, because of the very different types of data a log can usually contain. Maybe, we can put some basic fields into the baseclass. (i think every log has a field with a date, uh? ;)Perhaps a third and 4th parameters, giving either bytes or lines to skip, and the 4th parameter to give how many bytes or lines to read. So, to read the first 10 lines, you go with ::Read('foo', 'bar', 0, 10) and for lines 100-200, ::Read('foo', 'bar', 100, 100).So, for instance, to read your apache access log and count the total number of bytes served, you might do something like:$entries = Log_Parser::Read('/var/log/apache/access.log', 'apache_access');Thats what I thought about. There should be some mor optional parameters (like lines to parse, because a 75 MB apache-log could freeze your server for sometime... ;). Another point is the usage of multiple logfiles... But that should be no problem, imho.If you store the loglines in an array, it's easy to _merge them. Otherwise, filename could be a string for a single file, or an array of filenames. Whether you sort by timestamp, or leave them in the order they're read from the logfiles given, could be another issue.I'm not a great fan of accessors when a basic type will do just as well. But don't dump your method on my account - I think I'm fading into the minority in that regard.foreach ($entries as $e) { $total += $e->Element('size'); }I think an iterator would make this accesses more comfortable. I like such patterns and hopefully will implement a couple of them. But the way the elements would be stored will be the same. Think of this: while ($logline = $log->getLine('type')) { $total += $logline->getElement('size'); }I agree with your prposals. That would be some kind of cute approach. Would be great to have you in the team. I think with 2 or 3 people weNowhere near enough time, and I'm not interested in the idea enough to find the time for it. Sorry.Are there other logs, you'd like to include in the first developement-wave? I thinks it's ok to have a range of 3-5 different log-types for a general analysis and an initial release of the package.Apache is probably number 1, since PHP is a web language. Squid files would be useful to a lot of people, and various FTP daemons. - Matt