regexp fun. (issue XXVII)
| From: | Dieter Kneffel | Date: | Tue, 13 Jun 2000 19:08:02 +0000 |
| Subject: | regexp fun. (issue XXVII) | ||
| Groups: | php.general | ||
| Request: | Send a blank email to php-general+get-1744@lists.php.net to get a copy of this message | ||
here's a simple piece of script that magically links all URL looking
parts
= transforms a url in a html link '<a href=... >...</a>'
most of you should have seen this already on this list :-)
$text =
eregi_replace("(http|https|ftp)://([[:alnum:]/\n+-=%&:_.~?]+[#[:alnum:]+]*)",
"<a href=\"\\1://\\2\" target=\"_blank\">\\1://\\2</a>",
$text);
But here's the task:
Imagine you have a textfile, mixed with URLs and Links. so it can be
that there are
also html-links inside, e.g. with a different link text ( <a
href="http://this.is.a.long.domain.name.com/with/lot/of/characters/">more...</a>
so how can I modify the above regexp to only interpret URLs that are not
nested inside html anchors?
I already thought of checking for a blank/space in front of the
'http://xxx...' but this is only half way as some
URLs could be also at the very beginning of the text. So the only safe
solution would be to automatically link
only those URLs that are not quoted or at least have no quotes in front
of the 'h' ( bad: href="http://xxx.yy" |
good: http://xxx.yyy )
So, how can I extend the above regexp so it only modifies an URL that is
not nested in html???
Thanks,
dk