Re: Re: ucwords() vs title case

From: Date: Thu, 03 Jul 2014 13:45:30 +0000
Subject: Re: Re: ucwords() vs title case
References: 1 2 3 4 5 6 7 8  Groups: php.internals 
Request: Send a blank email to internals+get-75233@lists.php.net to get a copy of this message
On 3 July 2014 14:15, Tjerk Meesters <tjerk.meesters@gmail.com> wrote: > Hi! > > > On Thu, Jul 3, 2014 at 8:56 PM, Peter Cowburn <petercowburn@gmail.com> > wrote: > >> >> >> >> On 3 July 2014 13:39, Tjerk Meesters <tjerk.meesters@gmail.com> wrote: >> >>> On Wed, Jul 2, 2014 at 1:19 AM, Tjerk Meesters <tjerk.meesters@gmail.com >>> > >>> wrote: >>> >>> > Hi Kris, >>> > >>> > >>> > On Tue, Jul 1, 2014 at 7:25 AM, Kris Craig <kris.craig@gmail.com> >>> wrote: >>> > >>> >> On Mon, Jun 30, 2014 at 5:33 AM, Rowan Collins < >>> rowan.collins@gmail.com> >>> >> wrote: >>> >> >>> >> > Andrea Faulds wrote (on 30/06/2014): >>> >> > >>> >> >> On 30 Jun 2014, at 12:54, Tjerk Meesters <tjerk.meesters@gmail.com >>> > >>> >> >> wrote: >>> >> >> >>> >> >> Hi internals, >>> >> >>> >>> >> >>> I came across this old bug: >>> >> >>> https://bugs.php.net/bug.php?id=34407 >>> >> >>> >>> >> >>> >>> >> >>> >>> >> >>> Personally I find that the latter is too much of a departure from >>> >> what we >>> >> >>> currently have; a compromise could be to treat punctuation as a >>> word >>> >> >>> delimiter. >>> >> >>> >>> >> >> Hmm. Why not make it follow what \b in a regex would do, looking >>> for >>> >> >> “word boundaries”? >>> >> >> >>> >> > >>> >> > Unfortunately, the cleverer you try to be, the more edge cases you >>> find. >>> >> > For instance, using \b will capitalise the 's' after an >>> >> > apostrophe, >>> >> e.g. in >>> >> > "Andrea'S Suggestion". >>> >> > >>> >> > The function we have in our code base at the moment looks like this: >>> >> > >>> >> > function smart_uc_words($string) >>> >> > { >>> >> > $string = strtolower(trim($string)); >>> >> > // Capitalise any word char preceded by a non-word char >>> other >>> >> than >>> >> > an apostrophe >>> >> > $string = >>> >> > preg_replace_callback('/(?<!\w|\')(\w)/', >>> >> function($m){ >>> >> > return strtoupper($m[1]); }, $string); >>> >> > // Capitalise any word char which comes between an >>> apostrophe >>> >> and >>> >> > another word char >>> >> > $string = >>> >> > preg_replace_callback('/(?<=\')(\w)(?=\w)/', >>> >> > function($m){ return strtoupper($m[1]); }, $string); >>> >> > >>> >> > return $string; >>> >> > } >>> >> > >>> >> >>> >> What about leaving the default behavior as-is but adding an optional >>> >> argument to specify how to determine these boundaries? So if you did >>> >> something like ucwords( "hello, world!", '\b' ) or >>> >> ucwords( "hello, >>> >> world!", array( ' ', '.', ... ) ), the user could >>> >> control the behavior >>> >> while existing ucwords( $arg ) code would behave as it does now >>> without >>> >> any >>> >> BC. >>> >> >>> > >>> > Yeah, that seems like an option, so basically how ?ai >>> > :‹ܨҏÀ{¦trim() works too; >>> > treat these characters as word boundaries (default is " \t\r\n"). >>> > >>> > ucwords("hello (new) world", " ()"); >>> > >>> > I'll prepare a PR for this and see how far that takes us :) let me >>> know if >>> > you guys have any other ideas. >>> > >>> >>> I've created a PR here: >>> https://github.com/php/php-src/pull/706 >> >> >> Your previous mail mentioned, "so basically how trim() >> works too", but >> the PR doesn't quite do that. >> > > That's somewhat embarrassing; I didn't realise that character ranges are > supported in trim() =S > > Despite this oversight, I personally don't see a practical need in > supporting a character range because the given characters are not likely to > be letters, but rather hyphens, braces, punctuation marks, spaces, etc. I > was also hoping to keep the function rather simple :) > The charmasks aren't limited to letters only. You could go crazy and use "\0../;..@[..`{..\x7F" for everything non-alphanumeric in ASCII, if you really wanted to. That said, I have no particular preference one way or the other (for ucwords()) and was mostly just clarifying the point about working how trim() works, or not. > > >> >> Should ucwords() also accept character ranges, just like trim()? i.e., >> ucwords("Foo bar", "a..z"); [not a very practical example, I know] >> >> >>> >>> >>> If there are no objections I would like to commit this into 5.4 onwards >>> somewhere next week. >>> >>> Thanks. >>> >>> >>> > >>> > >>> > >>> >> --Kris >>> >> >>> > >>> > >>> > >>> > -- >>> > -- >>> > Tjerk >>> > >>> >>> >>> >>> -- >>> -- >>> Tjerk >>> >> >> > > > -- > -- > Tjerk >

« previous php.internals (#75233) next »