Re: Re: ucwords() vs title case
| From: | Tjerk Meesters | Date: | Tue, 01 Jul 2014 17:19:20 +0000 |
| Subject: | Re: Re: ucwords() vs title case | ||
| References: | 1 2 3 4 | Groups: | php.internals |
| Request: | Send a blank email to internals+get-75163@lists.php.net to get a copy of this message | ||
Hi Kris,
On Tue, Jul 1, 2014 at 7:25 AM, Kris Craig <kris.craig@gmail.com> wrote:
> On Mon, Jun 30, 2014 at 5:33 AM, Rowan Collins <rowan.collins@gmail.com>
> wrote:
>
> > Andrea Faulds wrote (on 30/06/2014):
> >
> >> On 30 Jun 2014, at 12:54, Tjerk Meesters <tjerk.meesters@gmail.com>
> >> wrote:
> >>
> >> Hi internals,
> >>>
> >>> I came across this old bug:
> >>> https://bugs.php.net/bug.php?id=34407
> >>>
> >>>
> >>>
> >>> Personally I find that the latter is too much of a departure from what
> we
> >>> currently have; a compromise could be to treat punctuation as a word
> >>> delimiter.
> >>>
> >> Hmm. Why not make it follow what \b in a regex would do, looking for
> >> “word boundaries”?
> >>
> >
> > Unfortunately, the cleverer you try to be, the more edge cases you find..
> > For instance, using \b will capitalise the 's' after an apostrophe, e.g..
> in
> > "Andrea'S Suggestion".
> >
> > The function we have in our code base at the moment looks like this:
> >
> > function smart_uc_words($string)
> > {
> > $string = strtolower(trim($string));
> > // Capitalise any word char preceded by a non-word char other
> than
> > an apostrophe
> > $string = preg_replace_callback('/(?<!\w|\')(\w)/',
> > function($m){
> > return strtoupper($m[1]); }, $string);
> > // Capitalise any word char which comes between an apostrophe and
> > another word char
> > $string = preg_replace_callback('/(?<=\')(\w)(?=\w)/',
> > function($m){ return strtoupper($m[1]); }, $string);
> >
> > return $string;
> > }
> >
>
> What about leaving the default behavior as-is but adding an optional
> argument to specify how to determine these boundaries? So if you did
> something like ucwords( "hello, world!", '\b' ) or ucwords( "hello,
> world!", array( ' ', '.', ... ) ), the user could control the behavior
> while existing ucwords( $arg ) code would behave as it does now without any
> BC.
>
Yeah, that seems like an option, so basically how
trim() works too; treat
these characters as word boundaries (default is " \t\r\n").
ucwords("hello (new) world", " ()");
I'll prepare a PR for this and see how far that takes us :) let me know if
you guys have any other ideas.
> --Kris
>
--
--
Tjerk