Re: Re: ucwords() vs title case

From: Date: Thu, 03 Jul 2014 13:15:41 +0000
Subject: Re: Re: ucwords() vs title case
References: 1 2 3 4 5 6 7  Groups: php.internals 
Request: Send a blank email to internals+get-75232@lists.php.net to get a copy of this message
Hi! On Thu, Jul 3, 2014 at 8:56 PM, Peter Cowburn <petercowburn@gmail.com> wrote: > > > > On 3 July 2014 13:39, Tjerk Meesters <tjerk.meesters@gmail.com> wrote: > >> On Wed, Jul 2, 2014 at 1:19 AM, Tjerk Meesters <tjerk.meesters@gmail.com> >> wrote: >> >> > Hi Kris, >> > >> > >> > On Tue, Jul 1, 2014 at 7:25 AM, Kris Craig <kris.craig@gmail.com> >> wrote: >> > >> >> On Mon, Jun 30, 2014 at 5:33 AM, Rowan Collins < >> rowan.collins@gmail.com> >> >> wrote: >> >> >> >> > Andrea Faulds wrote (on 30/06/2014): >> >> > >> >> >> On 30 Jun 2014, at 12:54, Tjerk Meesters <tjerk.meesters@gmail.com> >> >> >> wrote: >> >> >> >> >> >> Hi internals, >> >> >>> >> >> >>> I came across this old bug: >> >> >>> https://bugs.php.net/bug.php?id=34407 >> >> >>> >> >> >>> >> >> >>> >> >> >>> Personally I find that the latter is too much of a departure from >> >> what we >> >> >>> currently have; a compromise could be to treat punctuation as a >> word >> >> >>> delimiter. >> >> >>> >> >> >> Hmm. Why not make it follow what \b in a regex would do, looking for >> >> >> “word boundaries”? >> >> >> >> >> > >> >> > Unfortunately, the cleverer you try to be, the more edge cases you >> find. >> >> > For instance, using \b will capitalise the 's' after an apostrophe, >> >> e.g. in >> >> > "Andrea'S Suggestion". >> >> > >> >> > The function we have in our code base at the moment looks like this: >> >> > >> >> > function smart_uc_words($string) >> >> > { >> >> > $string = strtolower(trim($string)); >> >> > // Capitalise any word char preceded by a non-word char other >> >> than >> >> > an apostrophe >> >> > $string = preg_replace_callback('/(?<!\w|\')(\w)/', >> >> function($m){ >> >> > return strtoupper($m[1]); }, $string); >> >> > // Capitalise any word char which comes between an apostrophe >> >> and >> >> > another word char >> >> > $string = >> >> > preg_replace_callback('/(?<=\')(\w)(?=\w)/', >> >> > function($m){ return strtoupper($m[1]); }, $string); >> >> > >> >> > return $string; >> >> > } >> >> > >> >> >> >> What about leaving the default behavior as-is but adding an optional >> >> argument to specify how to determine these boundaries? So if you did >> >> something like ucwords( "hello, world!", '\b' ) or ucwords( >> >> "hello, >> >> world!", array( ' ', '.', ... ) ), the user could control >> >> the behavior >> >> while existing ucwords( $arg ) code would behave as it does now without >> >> any >> >> BC. >> >> >> > >> > Yeah, that seems like an option, so basically how >> > trim() works too; >> > treat these characters as word boundaries (default is " \t\r\n"). >> > >> > ucwords("hello (new) world", " ()"); >> > >> > I'll prepare a PR for this and see how far that takes us :) let me know >> if >> > you guys have any other ideas. >> > >> >> I've created a PR here: >> https://github.com/php/php-src/pull/706 > > > Your previous mail mentioned, "so basically how trim() > works too", but > the PR doesn't quite do that. > That's somewhat embarrassing; I didn't realise that character ranges are supported in trim() =S Despite this oversight, I personally don't see a practical need in supporting a character range because the given characters are not likely to be letters, but rather hyphens, braces, punctuation marks, spaces, etc. I was also hoping to keep the function rather simple :) > > Should ucwords() also accept character ranges, just like trim()? i.e., > ucwords("Foo bar", "a..z"); [not a very practical example, I know] > > >> >> >> If there are no objections I would like to commit this into 5.4 onwards >> somewhere next week. >> >> Thanks. >> >> >> > >> > >> > >> >> --Kris >> >> >> > >> > >> > >> > -- >> > -- >> > Tjerk >> > >> >> >> >> -- >> -- >> Tjerk >> > > -- -- Tjerk

« previous php.internals (#75232) next »