validate, produce complient URLs
| From: | bertrand Gugger | Date: | Mon, 13 Feb 2006 16:54:24 +0000 |
| Subject: | validate, produce complient URLs | ||
| Groups: | php.pear.dev | ||
| Request: | Send a blank email to pear-dev+get-41334@lists.php.net to get a copy of this message | ||
Bonjour,
It's a recurrent debate and I'm still stuck in what to do.
Everywhere we use URLs but nowhere we accord on what it should be to be compliant, widely acceptable.
That all thing is based on the weak rfc2396 [1] dated 1998 on which the only general accord is its weakness.
PEAR is naturally concerned by valid URLs as we use/produce them everywhere.
Once I proposed to rewrite Validate::uri() to bring it compliant [2] . Perhaps, as often, I had better shut my mouth.
We produce URLs in Text_Wiki [3] and want to be trusted on our output, so webmasters don't worry about it.
Net_URL is some central mean about it, gave some last discussion [4] proposing [5] as alternative.
... everywhere :) ...
Why is it important ?
* it's what engines every process in the Internet, as such it suffers from the guerrilla reigning, crackers use it to destroy your good will by using its weakness. Any script producing them from (even partial) user input *must* check their innocence.
* produced URLs are interpreted by user agents (browsers), we must comply to their process to do so independently of their settings, e.g., charset
* we eventually claim our output is compliant to some rules, say XHTML compliant.
So where is the problem ?
An URL if *full* is composed by several pieces (spaces are extra):
[scheme : ] [ // [userinfo @ ] [hostname | ip] [ : port] ] [ / path ] [ / ] [ ? query] [ # fragment]
yes, 7 parts.
As you can see ([]) every of these parts are optional. (I did not give the detail of each)
The rule is simple: if a part misses, then default is where we are.
(an example of nothing is an empty action attribute in some html <form> goes back to the same page / url)
Say, the url gives no scheme, then it takes it from the current document, "http:" if it's in a browser page...
It becomes hard when "no" host is given.
Luckily, we can have a first *single* "/" , then there starts the path, scheme and authority (userinfo+host+port) is extracted/completed from the current location.
Very often you don't have it, say "example.com" alone or :) ... "script.php"
When no starting "/", it's some "relative" url, meaning relative to the current doc (the one basically loaded, not necessarily the "current" file)
Anyway, the first (eventual) "?" starts the query part.
Here is the key: "absolute" or "relative" url !
As usual, php ensure what it does, "relative" are not proceeded by parse_url() and "//" is mandatory to start the host but "///" is allowed only in non standard "file" scheme (simple :) )
And pear does what it can, almost all url are "relative", and yes we must do, we can't escape as our script engine.
Hey !
That is already very long !
Let me finish by telling my tests based on [2], php's parse_url or [5] (no, I did not test Net_URL) are all good but never for everything, each has some "hole".
My last plan is to apply some "mekanik destruktiv kommando" inspired from [5]. Search a scheme first, then a question mark ... then subdivide/complete components.
For those thinking I'm crazy, so are URLs ... don't use them.
cheers
--
toggg
P.S. I've no idea what's the difference between URI and URL , I just want something practical,.
Purists: please open some other thread !
[1] http://www.faqs.org/rfcs/rfc2396.html
[2] http://cvs.php.net/viewcvs.cgi/pear/Validate/Validate.php
[3] http://pear.php.net/package/Text_Wiki
[4] http://beeblex.com/lists/index.php/php.pear.dev/41206?s=l%3Aphp.pear.dev+net_url
[5] http://baseclass.modulweb.dk/urlvalidator/viewsource.php