note 27501 deleted from function.ereg-replace by tularis
| From: | tularis@php.net | Date: | Sun, 29 Aug 2004 12:16:03 +0000 |
| Subject: | note 27501 deleted from function.ereg-replace by tularis | ||
| References: | 1 | Groups: | php.notes |
| Request: | Send a blank email to php-notes+get-75562@lists.php.net to get a copy of this message | ||
Note Submitter: xmontero at dsitelecom dot com
----
To scan for mails in an html documents, I've followed the "djworld at php dot net"
method, but I found several interesting issues that I'd like to share:
1) The range "A-Z" has a typo and is written with lower z instead of upper Z, so caution
if you copy/paste from there.
2) I don't know how his example works for multi-level subdomains, as he say, because the only
"dot" after the "@" sign is placed in the last block of 3 letters, so it works
for dude@domain.co.uk, but does not work for dude@long_department.long_company.com, so I added a
(dot + anything)* for the subdomain to be infinite-level.
3) Also the last limit {1,3} does accept "classical" top level domains like
".es", ".fr" or ".it" and the common ones: ".com/net/org",
but do not allow matching for the newer domains ended in ".info" or things like that. I
decided to increment that to 5 chars, but also keep in mind that if ever top level domains increase
in length, this is to be changed.
So, my expression ends up like this:
'[_a-zA-Z0-9\-]+' .
'(\.[_a-zA-Z0-9\-]+)*' .
'\@' .
'[_a-zA-Z0-9\-]+' .
'(\.[_a-zA-Z0-9\-]+)*' .
'(\.[a-zA-Z]{1,5})+'
Also, preg_replace works with that expression and is much faster. In fact, I use that preg instead
ereg.
Also, for those interested, to scan an HTML page for mails, I do this:
function txt_to_mails( $text )
{
// Scanning for the character @ to
// find emails.
set_time_limit( 30 );
// I remove italic, bold and underline tags, because
// they should "enfatize" only the user but not the
// domain or so and I don't wont to miss those emails.
$tmp = ereg_replace( '(<(/)*b>)|(<(/)*u>)|(<(/)*i>)', '',
$text ); // bold,under,italic
// I convert ANYTHING that cannot be a valid
// character for an email into a space, so I avoid
// rare things like http://somedomain.com/@hahaha.com
// to be scanned.
$tmp = preg_replace('([^_a-zA-Z0-9\@\-\.])', ' ', $text );
// I then split by the spaces, filling an array and
// making every entry in the array a set of characters
// valid to form an email.
// The following will produce an array with an item per 'word'
$alllines = explode( " ", $tmp );
// I then process line by line
reset( $alllines );
$selectedlines = array();
while( list( $key, $val ) = each( $alllines ) )
{
// If there is no @, discard the line
if( strstr( $val, "@" ) )
{
// I trim out the ending dots, so this:
// dude@company.com. becomes this:
// dude@company.com without the
// trailing dot.
$val = preg_replace( '((\.)+$)', '', $val );
// If the result is an email push into the result.
if( ereg(
'[_a-zA-Z0-9\-]+' .
'(\.[_a-zA-Z0-9\-]+)*' .
'\@' .
'[_a-zA-Z0-9\-]+' .
'(\.[_a-zA-Z0-9\-]+)*' .
'(\.[a-zA-Z]{1,5})+'
, $val ))
{
array_push( $selectedlines, $val );
}
}
}
return $selectedlines;
}
This can be used, for example, if you want to know if your employees are looking for jobs outside
your company monitoring the jobs with webs and comparing the emails found with the ones in your
company, or if your boss asks you to track 200 different web-based forums to see if there is
activity from the people he contracted to dynamize some products there.
See you and hope to help :-)
Xavier Montero.