Re: Re: ucwords() vs title case
| From: | Tjerk Meesters | Date: | Thu, 03 Jul 2014 12:39:02 +0000 |
| Subject: | Re: Re: ucwords() vs title case | ||
| References: | 1 2 3 4 5 | Groups: | php.internals |
| Request: | Send a blank email to internals+get-75230@lists.php.net to get a copy of this message | ||
On Wed, Jul 2, 2014 at 1:19 AM, Tjerk Meesters <tjerk.meesters@gmail.com>
wrote:
> Hi Kris,
>
>
> On Tue, Jul 1, 2014 at 7:25 AM, Kris Craig <kris.craig@gmail.com> wrote:
>
>> On Mon, Jun 30, 2014 at 5:33 AM, Rowan Collins <rowan.collins@gmail.com>
>> wrote:
>>
>> > Andrea Faulds wrote (on 30/06/2014):
>> >
>> >> On 30 Jun 2014, at 12:54, Tjerk Meesters <tjerk.meesters@gmail.com>
>> >> wrote:
>> >>
>> >> Hi internals,
>> >>>
>> >>> I came across this old bug:
>> >>> https://bugs.php.net/bug.php?id=34407
>> >>>
>> >>>
>> >>>
>> >>> Personally I find that the latter is too much of a departure from
>> what we
>> >>> currently have; a compromise could be to treat punctuation as a word
>> >>> delimiter.
>> >>>
>> >> Hmm. Why not make it follow what \b in a regex would do, looking for
>> >> “word boundaries”?
>> >>
>> >
>> > Unfortunately, the cleverer you try to be, the more edge cases you find.
>> > For instance, using \b will capitalise the 's' after an apostrophe,
>> e.g. in
>> > "Andrea'S Suggestion".
>> >
>> > The function we have in our code base at the moment looks like this:
>> >
>> > function smart_uc_words($string)
>> > {
>> > $string = strtolower(trim($string));
>> > // Capitalise any word char preceded by a non-word char other
>> than
>> > an apostrophe
>> > $string = preg_replace_callback('/(?<!\w|\')(\w)/',
>> function($m){
>> > return strtoupper($m[1]); }, $string);
>> > // Capitalise any word char which comes between an apostrophe
>> and
>> > another word char
>> > $string = preg_replace_callback('/(?<=\')(\w)(?=\w)/',
>> > function($m){ return strtoupper($m[1]); }, $string);
>> >
>> > return $string;
>> > }
>> >
>>
>> What about leaving the default behavior as-is but adding an optional
>> argument to specify how to determine these boundaries? So if you did
>> something like ucwords( "hello, world!", '\b' ) or ucwords(
>> "hello,
>> world!", array( ' ', '.', ... ) ), the user could control the
>> behavior
>> while existing ucwords( $arg ) code would behave as it does now without
>> any
>> BC.
>>
>
> Yeah, that seems like an option, so basically how
trim() works
> too;
> treat these characters as word boundaries (default is " \t\r\n").
>
> ucwords("hello (new) world", " ()");
>
> I'll prepare a PR for this and see how far that takes us :) let me know if
> you guys have any other ideas.
>
I've created a PR here: https://github.com/php/php-src/pull/706
If there are no objections I would like to commit this into 5.4 onwards
somewhere next week.
Thanks.
>
>
>
>> --Kris
>>
>
>
>
> --
> --
> Tjerk
>
--
--
Tjerk