Re: Re: ucwords() vs title case
| From: | Tjerk Meesters | Date: | Thu, 03 Jul 2014 13:15:41 +0000 |
| Subject: | Re: Re: ucwords() vs title case | ||
| References: | 1 2 3 4 5 6 7 | Groups: | php.internals |
| Request: | Send a blank email to internals+get-75232@lists.php.net to get a copy of this message | ||
Hi!
On Thu, Jul 3, 2014 at 8:56 PM, Peter Cowburn <petercowburn@gmail.com>
wrote:
>
>
>
> On 3 July 2014 13:39, Tjerk Meesters <tjerk.meesters@gmail.com> wrote:
>
>> On Wed, Jul 2, 2014 at 1:19 AM, Tjerk Meesters <tjerk.meesters@gmail.com>
>> wrote:
>>
>> > Hi Kris,
>> >
>> >
>> > On Tue, Jul 1, 2014 at 7:25 AM, Kris Craig <kris.craig@gmail.com>
>> wrote:
>> >
>> >> On Mon, Jun 30, 2014 at 5:33 AM, Rowan Collins <
>> rowan.collins@gmail.com>
>> >> wrote:
>> >>
>> >> > Andrea Faulds wrote (on 30/06/2014):
>> >> >
>> >> >> On 30 Jun 2014, at 12:54, Tjerk Meesters <tjerk.meesters@gmail.com>
>> >> >> wrote:
>> >> >>
>> >> >> Hi internals,
>> >> >>>
>> >> >>> I came across this old bug:
>> >> >>> https://bugs.php.net/bug.php?id=34407
>> >> >>>
>> >> >>>
>> >> >>>
>> >> >>> Personally I find that the latter is too much of a departure from
>> >> what we
>> >> >>> currently have; a compromise could be to treat punctuation as a
>> word
>> >> >>> delimiter.
>> >> >>>
>> >> >> Hmm. Why not make it follow what \b in a regex would do, looking for
>> >> >> “word boundaries”?
>> >> >>
>> >> >
>> >> > Unfortunately, the cleverer you try to be, the more edge cases you
>> find.
>> >> > For instance, using \b will capitalise the 's' after an apostrophe,
>> >> e.g. in
>> >> > "Andrea'S Suggestion".
>> >> >
>> >> > The function we have in our code base at the moment looks like this:
>> >> >
>> >> > function smart_uc_words($string)
>> >> > {
>> >> > $string = strtolower(trim($string));
>> >> > // Capitalise any word char preceded by a non-word char other
>> >> than
>> >> > an apostrophe
>> >> > $string = preg_replace_callback('/(?<!\w|\')(\w)/',
>> >> function($m){
>> >> > return strtoupper($m[1]); }, $string);
>> >> > // Capitalise any word char which comes between an apostrophe
>> >> and
>> >> > another word char
>> >> > $string =
>> >> > preg_replace_callback('/(?<=\')(\w)(?=\w)/',
>> >> > function($m){ return strtoupper($m[1]); }, $string);
>> >> >
>> >> > return $string;
>> >> > }
>> >> >
>> >>
>> >> What about leaving the default behavior as-is but adding an optional
>> >> argument to specify how to determine these boundaries? So if you did
>> >> something like ucwords( "hello, world!", '\b' ) or ucwords(
>> >> "hello,
>> >> world!", array( ' ', '.', ... ) ), the user could control
>> >> the behavior
>> >> while existing ucwords( $arg ) code would behave as it does now without
>> >> any
>> >> BC.
>> >>
>> >
>> > Yeah, that seems like an option, so basically how
>> >
trim() works too;
>> > treat these characters as word boundaries (default is " \t\r\n").
>> >
>> > ucwords("hello (new) world", " ()");
>> >
>> > I'll prepare a PR for this and see how far that takes us :) let me know
>> if
>> > you guys have any other ideas.
>> >
>>
>> I've created a PR here:
>> https://github.com/php/php-src/pull/706
>
>
> Your previous mail mentioned, "so basically how trim()
> works too", but
> the PR doesn't quite do that.
>
That's somewhat embarrassing; I didn't realise that character ranges are
supported in trim() =S
Despite this oversight, I personally don't see a practical need in
supporting a character range because the given characters are not likely to
be letters, but rather hyphens, braces, punctuation marks, spaces, etc. I
was also hoping to keep the function rather simple :)
>
> Should ucwords() also accept character ranges, just like trim()? i.e.,
> ucwords("Foo bar", "a..z"); [not a very practical example, I know]
>
>
>>
>>
>> If there are no objections I would like to commit this into 5.4 onwards
>> somewhere next week.
>>
>> Thanks.
>>
>>
>> >
>> >
>> >
>> >> --Kris
>> >>
>> >
>> >
>> >
>> > --
>> > --
>> > Tjerk
>> >
>>
>>
>>
>> --
>> --
>> Tjerk
>>
>
>
--
--
Tjerk