Re: Re: ucwords() vs title case
| From: | Peter Cowburn | Date: | Thu, 03 Jul 2014 13:45:30 +0000 |
| Subject: | Re: Re: ucwords() vs title case | ||
| References: | 1 2 3 4 5 6 7 8 | Groups: | php.internals |
| Request: | Send a blank email to internals+get-75233@lists.php.net to get a copy of this message | ||
On 3 July 2014 14:15, Tjerk Meesters <tjerk.meesters@gmail.com> wrote:
> Hi!
>
>
> On Thu, Jul 3, 2014 at 8:56 PM, Peter Cowburn <petercowburn@gmail.com>
> wrote:
>
>>
>>
>>
>> On 3 July 2014 13:39, Tjerk Meesters <tjerk.meesters@gmail.com> wrote:
>>
>>> On Wed, Jul 2, 2014 at 1:19 AM, Tjerk Meesters <tjerk.meesters@gmail.com
>>> >
>>> wrote:
>>>
>>> > Hi Kris,
>>> >
>>> >
>>> > On Tue, Jul 1, 2014 at 7:25 AM, Kris Craig <kris.craig@gmail.com>
>>> wrote:
>>> >
>>> >> On Mon, Jun 30, 2014 at 5:33 AM, Rowan Collins <
>>> rowan.collins@gmail.com>
>>> >> wrote:
>>> >>
>>> >> > Andrea Faulds wrote (on 30/06/2014):
>>> >> >
>>> >> >> On 30 Jun 2014, at 12:54, Tjerk Meesters <tjerk.meesters@gmail.com
>>> >
>>> >> >> wrote:
>>> >> >>
>>> >> >> Hi internals,
>>> >> >>>
>>> >> >>> I came across this old bug:
>>> >> >>> https://bugs.php.net/bug.php?id=34407
>>> >> >>>
>>> >> >>>
>>> >> >>>
>>> >> >>> Personally I find that the latter is too much of a departure from
>>> >> what we
>>> >> >>> currently have; a compromise could be to treat punctuation as a
>>> word
>>> >> >>> delimiter.
>>> >> >>>
>>> >> >> Hmm. Why not make it follow what \b in a regex would do, looking
>>> for
>>> >> >> “word boundaries”?
>>> >> >>
>>> >> >
>>> >> > Unfortunately, the cleverer you try to be, the more edge cases you
>>> find.
>>> >> > For instance, using \b will capitalise the 's' after an
>>> >> > apostrophe,
>>> >> e.g. in
>>> >> > "Andrea'S Suggestion".
>>> >> >
>>> >> > The function we have in our code base at the moment looks like this:
>>> >> >
>>> >> > function smart_uc_words($string)
>>> >> > {
>>> >> > $string = strtolower(trim($string));
>>> >> > // Capitalise any word char preceded by a non-word char
>>> other
>>> >> than
>>> >> > an apostrophe
>>> >> > $string =
>>> >> > preg_replace_callback('/(?<!\w|\')(\w)/',
>>> >> function($m){
>>> >> > return strtoupper($m[1]); }, $string);
>>> >> > // Capitalise any word char which comes between an
>>> apostrophe
>>> >> and
>>> >> > another word char
>>> >> > $string =
>>> >> > preg_replace_callback('/(?<=\')(\w)(?=\w)/',
>>> >> > function($m){ return strtoupper($m[1]); }, $string);
>>> >> >
>>> >> > return $string;
>>> >> > }
>>> >> >
>>> >>
>>> >> What about leaving the default behavior as-is but adding an optional
>>> >> argument to specify how to determine these boundaries? So if you did
>>> >> something like ucwords( "hello, world!", '\b' ) or
>>> >> ucwords( "hello,
>>> >> world!", array( ' ', '.', ... ) ), the user could
>>> >> control the behavior
>>> >> while existing ucwords( $arg ) code would behave as it does now
>>> without
>>> >> any
>>> >> BC.
>>> >>
>>> >
>>> > Yeah, that seems like an option, so basically how ?ai
>>> > :‹ܨҏÀ{¦trim() works too;
>>> > treat these characters as word boundaries (default is " \t\r\n").
>>> >
>>> > ucwords("hello (new) world", " ()");
>>> >
>>> > I'll prepare a PR for this and see how far that takes us :) let me
>>> know if
>>> > you guys have any other ideas.
>>> >
>>>
>>> I've created a PR here:
>>> https://github.com/php/php-src/pull/706
>>
>>
>> Your previous mail mentioned, "so basically how
trim()
>> works too", but
>> the PR doesn't quite do that.
>>
>
> That's somewhat embarrassing; I didn't realise that character ranges are
> supported in trim() =S
>
> Despite this oversight, I personally don't see a practical need in
> supporting a character range because the given characters are not likely to
> be letters, but rather hyphens, braces, punctuation marks, spaces, etc. I
> was also hoping to keep the function rather simple :)
>
The charmasks aren't limited to letters only. You could go crazy and use
"\0../;..@[..`{..\x7F" for everything non-alphanumeric in ASCII, if you
really wanted to.
That said, I have no particular preference one way or the other (for
ucwords()) and was mostly just clarifying the point about working how
trim() works, or not.
>
>
>>
>> Should ucwords() also accept character ranges, just like trim()? i.e.,
>> ucwords("Foo bar", "a..z"); [not a very practical example, I know]
>>
>>
>>>
>>>
>>> If there are no objections I would like to commit this into 5.4 onwards
>>> somewhere next week.
>>>
>>> Thanks.
>>>
>>>
>>> >
>>> >
>>> >
>>> >> --Kris
>>> >>
>>> >
>>> >
>>> >
>>> > --
>>> > --
>>> > Tjerk
>>> >
>>>
>>>
>>>
>>> --
>>> --
>>> Tjerk
>>>
>>
>>
>
>
> --
> --
> Tjerk
>