Re: finally: new authors

From: Date: Sun, 04 May 2003 11:26:39 +0000
Subject: Re: finally: new authors
References: 1  Groups: php.doc 
Request: Send a blank email to phpdoc+get-969353146@lists.php.net to get a copy of this message
BTW: I hate politics. Since I have no stake, it's easier. On Sunday, May 4, 2003, at 02:59 AM, Gabor Hojtsy wrote:
If someone added or edited 50% of the pages, but hasn't bothered later fixing typos, does that devalue their past work as compared to recent edits? Has the later work now surpassed their original work? What is a fair metric for determining this?
I don't think that future editing devalues someones work. But I think it can be undestood that we would like to give more credit to those who actually actively work on something.
But how do we determine this? Past contribution vs. future?
I am struck that all manual contributors, for the most part, are professional coders, who have access to *all* of the data needed to determine the actual level of each human, individual, contribution,
Do we? That would make me happy :)
I should think so.
-Let's
    say for someone to become a member of the offical PHP
    Documentation Group (ie. listed on the frontpage of the
    manual), he needs at least three positive votes from listed
    authors, and no negative ones.
Contribution listing as a high-school popularity contest, with the "cool people" doing the voting??? This strikes me as a possible disincentive, politicizing a task.
Well, as you see, we are going to add 'many' people to the credits list either we go on the human route or the computer automated one.
What ensures that the first list will not be the last (for a long time) one?
A long time ago, I wrote a bit about this. The notes editors can (and should) easily fold the *value* of a note without folding in the direct note, verbatim... the text author is then the note editor, even if the *idea* to add/change came from someone else..... back to the listing issue, though.
Well, to be fair most of the time, notes editors are simply copying over the user note with slight editing.
Bad form, this exposes the docs to copyright lawsuits. :-(
We should not expect them IMHO to rewrite the text just because of legality questions. We should put the notes into the legal system as 'free to be used by the docteam' IMHO.
Adding notes, in that case, should notify the adder that they are releasing *all* intellectual property, copyright, etc etc. Or, we should add each notes user to the list of manual authors.
Might I propose a metaphor: In american movies, the biggest star people are before the movie, the rest of the people are afterwards. So, why not a "major credit" listing on the main credit page or front page, and "the tons of others" on another page?
And do anybody read the list at the end of the film? Not in Hungary.
*Some* do here.
The lights are turned on as that list starts, and people go out. Those lists are also cut in TV broadcasting and replaced with commercials. So do having a list of 400 contributors will be a real show of respect? I don't think so.
Will a list of 400 be treated with respect on the front page? The credits page? This is the issue.
If you convince _people_, rather than _computers_ that you deserve to be on the credits list, the list will be a whole lot more meaningful, as Goba mentioned.
I disagree that this is a good metric, and think it is a poor metric, as well. It provides direct incentive for people to be well liked first and foremost, rather than accurate, or frequently contributing.
Well, even if this does not seem to be a good metric, this seems to be much better to me, than any automation, and not because I won't get onto the list if an automated system is used, or because I can't get my mates to be on the list. But because I think to give real credit, and for the contributors to feel that they are respected, they should not be picked by a computer.
A bad computer system and a good computer system do not mean the computers are bad or good. It means the programmers are bad or good.
I feel much better when receiving a mail from any of my friends, though technically it is the same thing. But one mail reflects the feelings and care of someone, the other reflects some program's automated behaviour.
How would you feel about an automated system that added points for your peers appreciating you?
Something possibly worth discussing: CVS doesn't just track number of commits, it also tracks lines changed, what lines were changed, and *how much change was involved in a file*, with a little coaxing. Let's take a few examples: Author A: Changes one punctuation character a day. 365 commits per year. However, comparing the old lines versus new lines only shows 365 individual characters added/changed/removed per year. They are well liked by others. Author B: Changes a page once a month, rewriting whole sections. Only 12 commits per year. However, comparing old lines to new lines show 30,000 characters added/changed/removed per year. They are almost universally disliked, because they are constantly changing the work of others, and hold unpopular social and political views. Author C: Changes 5 pages every day. Some pages are one character, some are entirely new paragraphs/examples/etc. 1825 commits, 350,000 added/changed/removed characters per year. They are viewed with some suspicion by many others. Depending on the system used, A or B or C may be a 'top listed author". Here is what *I* would suggest: 1) Author listings are based on monthly added/changed character counts, and changed/updated monthly. 2) Major authors get their own listing, minor authors are separate. Can this system be "gamed"? Yes, but I think the gaming of character counts is a lot more obvious than the gaming of people.
Well, I suppose you have not seen some of my last commits to 'phpweb'. I have a newly installed linux system (upgraded yesterday), and I had the 'remove whitespace from the end of lines' option checked in by default in my IDE. Therefore I have committed some files with many chars changed, although I have done small real changes to the files.
All long whitespace is ignorable. As are character returns. I do not think they should be counted, and CVS counting them is annoying to me, but easily fixable. :-)
Let me have another more closer example. In phpdoc we have a kind of 'wrapping policy', basically lines should not be long. What if I add 12 chars into some paragraph? I need to rewrap it to make it readable. That means some words will go to new lines, and depending on the length of the para, many chars will be changed (let's say 15 lines), just because I have added one or two words.
If a filter bypasses those, and concatenates each file, lines changes don't matter.
What if someone commits a translated file by error? What if someone else fixes the problem? That would mean hundreds of changed lines for both of the guys. These are real examples, these are happening every week here!
100 lines of CR->CRLF are easily ignored, as are any other line changes. They are predictable.
What if I commit one file with windows newlines? The one guy who will correct me, will have another hundred lines. What if I commit one file with tabs instead of spaces? The guy who corrects me will have several more chars (1 tab = 4 spaces!).
That's obvious gaming and mistakes, and can be filtered out of "valued commits", and any half-sane filter would strip out cr/lf combinations before character comparison. The US comment is that "this isn't rocket science", meaning that we don't need 20 years of study, simple filters will do the job. Strip excess whitespace, and line characters, and the job becomes simple.
The automation will think that I am a cool guy, who added a great deal of content, or corrected many errors in the documentation, however this is not the case.
Well, that's *bad* automation. This is not 1986, we can now all afford computers which can compare thousands of lines in a short period of time.
CVS gaming is much more transparent Well, ok. But you are talking about using raw numbers. If you consider all the things above, and are convinced that using raw numbers won't show the quality of the work someone done, then you end up putting some guys looking behind the numbers. And then we get to the system I am proposing, decisionmaking by humans.
Humans, after dealing with their decision making vs. computers, for 19 years, have biases (bugs) too. I would feel much better about endorsing the human system if it could be over-ridden by numbers. If the numbers are wrong, fix the numeric system, as recoding computers seems much faster than re-coding humans. :-)
Even if noone is going to get on committing things with large change counts before anyone else can, these factors will make the numbers untrustable.
Only in a bad system. Do you have examples of things that cannot be easily programmed around?
than buying people at a conference beer all night long. It seems our points of view are quite different.
I have seen systems of popularity easily gamed, so I have some concerns.
I think I can't be convinced with a beer. I don't like beer anyway ;) [By the way, I would like to buy a laptop, in case anyone is interested ;)]
ROTFL.... what make/model? :-) -Bop Ronald Chmara Ronin Professional Consulting LLC 520-326-6109 "It can only be attributable to human error." --Hal.

« previous php.doc (#969353146) next »