BBO Discussion Forums: Rating systems (again?) - BBO Discussion Forums

Jump to content

  • 3 Pages +
  • 1
  • 2
  • 3
  • You cannot start a new topic
  • You cannot reply to this topic

Rating systems (again?)

#21 User is offline   helene_t 

  • The Abbess
  • PipPipPipPipPipPipPipPipPipPipPip
  • Group: Advanced Members
  • Posts: 17,401
  • Joined: 2004-April-22
  • Gender:Female
  • Location:Odense, Denmark
  • Interests:History, languages

Posted 2007-July-02, 14:52

fred, on Jul 2 2007, 06:29 PM, said:

I think that for every one BBO member who genuinely wants to track their progress and see if they are improving, there are many BBO members who do not want to know how poorly they really play.

I know perfectly well how bad I am, it's just that I don't want to be reminded of it constantly. :)

(Not going to repeat what has been said thousand times about the devastating social impact of rating systems).
The world would be such a happy place, if only everyone played Acol :) --- TramTicket
0

#22 User is offline   fred 

  • PipPipPipPipPipPipPip
  • Group: Advanced Members
  • Posts: 4,614
  • Joined: 2003-February-11
  • Gender:Male
  • Location:Las Vegas, USA

Posted 2007-July-02, 16:41

fred, on Jul 2 2007, 04:29 PM, said:

I think that for every one BBO member who genuinely wants to track their progress and see if they are improving, there are many BBO members who do not want to know how poorly they really play.

Please note that I was not trying to say that masses of BBO members play "poorly" in any absolute sense.

The type of rating systems used for online bridge do not attempt to measure absolute skill. Instead they measure skill relative to other players.

All I was trying to say was bridge players (and not just BBO members) tend to overestimate how well they play compared to other players.

Fred Gitelman
Bridge Base Inc.
www.bridgebase.com
0

#23 User is offline   mycroft 

  • Secretary Bird
  • PipPipPipPipPipPipPipPipPip
  • Group: Advanced Members
  • Posts: 8,414
  • Joined: 2003-July-12
  • Gender:Male
  • Location:Calgary, D18; Chapala, D16

Posted 2007-July-02, 16:56

Who was it who ran that poll of bridge players, and found that 90% of them were better than their partners?

Michael (who still thinks that "experts play with me" is the best rating I will get).
"Which is harder to find - a paranormal field agent, or someone competent who likes talking on the phone?"
"...You may return to your desk." "Thank you." -- Serena vs. Mr. Arthur, "Paranormal Helpline", EGS:NP
0

#24 User is offline   kgr 

  • PipPipPipPipPipPipPip
  • Group: Advanced Members
  • Posts: 3,457
  • Joined: 2003-April-11

Posted 2007-July-03, 03:47

I play most of the time in the main bridge club with pick-up partners (I hope Fred or his partners also do this anonimiously, to see how this works). The self-rating of a lot of players is really painfull. Even adding comments on players doesn't help. Some days ago I was not the table host and an expert was allowed to play with me. I had marked him with bad and that prooved to be true again. No surprise you have a lot of table hopping (..is this English?) in the main bridge club. (Some time ago a post of a new BBO player showed that this is a real issue).
I vote for having a peer rating system like on Ebay. This could have two ratings: Play and friendlyness. You could only receive on rate per player. Play rating could be wheighted (giving more weight to players that are rated high themselve; some weight to the numbers of boards that are played).
I think that a rating system like that would improve the friendliness and quality of the site: For a good rating you need to be friendly and play long enough to convince your partner of your good play.

Regards,
Koen
0

#25 User is offline   hotShot 

  • Axxx Axx Axx Axx
  • PipPipPipPipPipPipPip
  • Group: Advanced Members
  • Posts: 2,976
  • Joined: 2003-August-31
  • Gender:Male

Posted 2007-July-03, 05:34

A fair rating is bridge is much more complex than e.g in chess.

This is because:
- in one table bridge there is no true winner
- the best player don't have to get best score
- the "first" boards of any new partnership are below standard because of insufficient agreements
- any tournament result is a function of the field, so an average score is meaningless if the field does not stay the same
- you need a significant number of boards to make sure skill is more important than luck
- there is no undisputed way to transform a pair/team result into an individual rating

Ratings based on boards played in MBC with partners or opps changing every 2nd board are worthless.
Ratings based on tourneys with 3-6 boards played are worthless.
Ratings based of tourney results with 12-15 boards played and more than 30 tables are worthless, because you play so few opps, that the field each pair has is almost completely different.

So most boards played online are of no valuable use far a rating system.

Maybe team games on BBO could be used to get some sort of rating, but what about those players you rarely play team?
0

#26 User is offline   dosxtres 

  • PipPipPip
  • Group: Full Members
  • Posts: 92
  • Joined: 2005-September-05

Posted 2007-July-03, 07:51

I think the implementation of a rating system is a very big problem.
A good rating system is almost impossible.
It depends where do u play, who do you play with, what is the field...and more.

Because wc want to launch games easyly with the best players, but maybe the dont know all other. Experts want to play with experts and wc. Advaced want to play with adv and experts. Int want to play with adv. Beg want to.....

That is how they can probe that they are learning, they enjoy the game, and that is like in all aspects of the life, things have been doing. Ones must learn from others who know more.

Anyway, if something could be done, i think it could be with something that anyone can make visible to people in his profile. Like clubs you are joined, and an status that could be managed by each club stuff.

say, Expert Mountain club. Someone ask for joining.
At the end of the trial period, the Expert Mountain club reject him or accept.
So, they give him, say, the expert status (by this club point of view)

So, you can enable to people: Expert Mountain Club skill level: Expert

The rest is, you can trust some clubs or not, but you can have a clue.
And everything is up to the single people to enable this status.
0

#27 User is offline   helene_t 

  • The Abbess
  • PipPipPipPipPipPipPipPipPipPipPip
  • Group: Advanced Members
  • Posts: 17,401
  • Joined: 2004-April-22
  • Gender:Female
  • Location:Odense, Denmark
  • Interests:History, languages

Posted 2007-July-04, 00:30

I wonder what sick soul came up with the idea of computing individual ratings from online selected-partner-game results. One can argue about the meaningfulness of such a scheme but the averse social implications must be obvious to everyone, even in advance. Much more now that we have learned from the experience from other sites.

I would have more sympathy for individual ratings based on one-human-plus-7-GIBs team matches.
The world would be such a happy place, if only everyone played Acol :) --- TramTicket
0

#28 User is offline   nige1 

  • 5-level belongs to me
  • PipPipPipPipPipPipPipPipPip
  • Group: Advanced Members
  • Posts: 9,128
  • Joined: 2004-August-30
  • Gender:Male
  • Location:Glasgow Scotland
  • Interests:Poems Computers

  Posted 2007-July-04, 19:34

helene_t, on Jul 4 2007, 01:30 AM, said:

I wonder what sick soul came up with the idea of computing individual ratings from online selected-partner-game results. One can argue about the meaningfulness of such a scheme but the averse social implications must be obvious to everyone, even in advance. Much more now that we have learned from the experience from other sites.

I would have more sympathy for individual ratings based on one-human-plus-7-GIBs team matches.

Most games are competitive by nature. Bridge certainly is. Unsurprisingly, at face-to-face bridge, the various master-point and gold-point rating schemes are among the most popular facilities provided by National Bridge Organizations.

On-line bridge allows more accurate computation of ratings, including quite accurate estimates of current form, if you play enough.

I see little harm in giving players feedback of this nature if they opt for it. The site can reveal such information only to individuals who are interested. Nobody else need be nauseated.
0

#29 User is offline   Al_U_Card 

  • PipPipPipPipPipPipPipPip
  • Group: Advanced Members
  • Posts: 6,080
  • Joined: 2005-May-16
  • Gender:Male

Posted 2007-July-05, 08:45

If you go to myhands, type in the player nick and ask for the last month's results, you get a good indication of his performance in terms of average imps and average mp % over all the hands played.

This is an indicator and might save the hand or two that you usually need to "spot the loony" and excuse yourself from the table. Would it lead people to attempt to "pump up" their numbers? Remains to be seen.
The Grand Design, reflected in the face of Chaos...it's a fluke!
0

#30 User is offline   hrothgar 

  • PipPipPipPipPipPipPipPipPipPipPip
  • Group: Advanced Members
  • Posts: 15,725
  • Joined: 2003-February-13
  • Gender:Male
  • Location:Natick, MA
  • Interests:Travel
    Cooking
    Brewing
    Hiking

Posted 2007-July-05, 09:23

nige1, on Jul 5 2007, 04:34 AM, said:

On-line bridge allows more accurate computation of ratings, including quite accurate estimates of current form, if you play enough.

Developing an accurate rating system is far from trivial. I suspect that a fair amount of engineering effort would be necessary to devise an accurate system.

This ignores the enormous ongoing effort that would be required to administer a rating system. As I've noted in the past, I don't see any reasonable way to develop a rating system that is both

1. Accurate
2. Can be understood by laymen

I recall all the bullshit necessary to try to explain the implementation of the Lehman system to average players. There was a never-endering queue of folks trying to understand how the ratings where calculated, arguing about implementation details, and trying to explain how the system was flawed because they're really much better than their score suggests...

Please note: This cropped up all the time with the Lehman system, which used a very simple algorithm. Personally, I suspect that an accurate rating is going to need to use a Kalman filter or some other similar approaches taken from signal processing. You are NEVER going to be able to explain the implementation to the average. Accordingly, the political issues will be much worse.

Personally, I'd prefer if Fred and Uday focused on more serious issues. Equally significant, its far from clear whether Fred / Uday have the technical expertise to address this type of project.

Here's my suggestion:

Bridge Browser provides all the raw data that folks would need to develop and test a rating system. If you think that you can develop an accurate system, go out and do so. Demonstrate that your system has good predictive power. Once you're done, come back and tell us about it. Folks can then debate whether or not this should be integrated into BBO. Conversely, if you aren't willing/able to do the necessary work, then stop talking about what can be done...

(BTW, I'll make my usual suggestion that you will probably have better luck if you initially focus on a system that is able to accurately rate partnerships rather than rating players)
Alderaan delenda est
0

#31 User is offline   awm 

  • PipPipPipPipPipPipPipPipPip
  • Group: Advanced Members
  • Posts: 8,707
  • Joined: 2005-February-09
  • Gender:Male
  • Location:Zurich, Switzerland

Posted 2007-July-05, 10:17

Trying to rate individuals based on information about performance of pairs seems to be a hard mathematical problem. It's made especially difficult because pairs can be more or less than the sum of the parts (depending on partnership experience and agreements).

However, on BBO we really have more information than just the net results of contracts. We could analyze each card played. In particular, I like the idea of computing double-dummy error rate. We could define:

(1) A play is a double-dummy error if it reduces the double-dummy trick total for the person making the play. It's easy to notice these when kibitzing using GIB.

(2) A play is a double-dummy contract-costing error if it changes the result of the hand from making to down (for declarer) or vice versa (for defense).

In principle we could compute these error rates. This has a number of nice features; in particular the error rates for declarer are mostly independent of partner. There are obviously a few things to watch out for, in particular:

(1) Everyone will have non-zero error rates, because double-dummy play often requires anti-percentage lines.

(2) Safety plays will often score as a double-dummy error, but not as a contract-costing error. Especially at IMP scoring it may be good to be aware of this. An alternative might be to weight the errors by IMP cost (so losing an overtrick costs only 1, but losing the contract costs 10 or whatever for a game).

(3) One can avoid contract-costing errors by constantly underbidding, so every contract rates to make +1 or +2.

(4) On defense, there will be impact from partner's signals/failure to signal to the degree that some mistakes will not be a defender's "fault."

However, despite the problems, this seems to me like a better way to measure skill than trying to look simply at end results.
Adam W. Meyerson
a.k.a. Appeal Without Merit
0

#32 User is offline   hotShot 

  • Axxx Axx Axx Axx
  • PipPipPipPipPipPipPip
  • Group: Advanced Members
  • Posts: 2,976
  • Joined: 2003-August-31
  • Gender:Male

Posted 2007-July-05, 10:30

awm, on Jul 5 2007, 06:17 PM, said:

However, on BBO we really have more information than just the net results of contracts. We could analyze each card played. In particular, I like the idea of computing double-dummy error rate. We could define:

(1) A play is a double-dummy error if it reduces the double-dummy trick total for the person making the play. It's easy to notice these when kibitzing using GIB.

(2) A play is a double-dummy contract-costing error if it changes the result of the hand from making to down (for declarer) or vice versa (for defense).

In principle we could compute these error rates. This has a number of nice features; in particular the error rates for declarer are mostly independent of partner. There are obviously a few things to watch out for, in particular:

I have done that and posted some results of that a while ago. But the system is limited to play and ignores bidding completely.
It is less significant as one might think. WC's and gib produce error rates of about 0.4 errors per board. Intermediates make about 0.8 and beginners reach 0.9+.
I had a few pickup partner with rates larger than 1, but i did not play enough boards with them to make the results mean something in a statistic.
0

#33 User is offline   junyi_zhu 

  • PipPipPipPipPip
  • Group: Full Members
  • Posts: 536
  • Joined: 2003-May-28
  • Location:Saltlake City

Posted 2007-July-06, 09:03

hotShot, on Jul 5 2007, 04:30 PM, said:

awm, on Jul 5 2007, 06:17 PM, said:

However, on BBO we really have more information than just the net results of contracts. We could analyze each card played. In particular, I like the idea of computing double-dummy error rate. We could define:

(1) A play is a double-dummy error if it reduces the double-dummy trick total for the person making the play. It's easy to notice these when kibitzing using GIB.

(2) A play is a double-dummy contract-costing error if it changes the result of the hand from making to down (for declarer) or vice versa (for defense).

In principle we could compute these error rates. This has a number of nice features; in particular the error rates for declarer are mostly independent of partner. There are obviously a few things to watch out for, in particular:

I have done that and posted some results of that a while ago. But the system is limited to play and ignores bidding completely.
It is less significant as one might think. WC's and gib produce error rates of about 0.4 errors per board. Intermediates make about 0.8 and beginners reach 0.9+.
I had a few pickup partner with rates larger than 1, but i did not play enough boards with them to make the results mean something in a statistic.

World Class and GIB make 0.4 errors per hand in card plays? That's just an insult of world class players, hahaha. GIB is generally weaker than intermediate players, not even in bidding, but also in card plays. It may make some tough double dummy contracts, however, it blows way more tricks in simple situations and it blows way more tricks in redoubled contracts than you can imagine.
0

#34 User is offline   junyi_zhu 

  • PipPipPipPipPip
  • Group: Full Members
  • Posts: 536
  • Joined: 2003-May-28
  • Location:Saltlake City

Posted 2007-July-06, 09:11

awm, on Jul 5 2007, 04:17 PM, said:

Trying to rate individuals based on information about performance of pairs seems to be a hard mathematical problem. It's made especially difficult because pairs can be more or less than the sum of the parts (depending on partnership experience and agreements).

However, on BBO we really have more information than just the net results of contracts. We could analyze each card played. In particular, I like the idea of computing double-dummy error rate. We could define:

(1) A play is a double-dummy error if it reduces the double-dummy trick total for the person making the play. It's easy to notice these when kibitzing using GIB.

(2) A play is a double-dummy contract-costing error if it changes the result of the hand from making to down (for declarer) or vice versa (for defense).

In principle we could compute these error rates. This has a number of nice features; in particular the error rates for declarer are mostly independent of partner. There are obviously a few things to watch out for, in particular:

(1) Everyone will have non-zero error rates, because double-dummy play often requires anti-percentage lines.

(2) Safety plays will often score as a double-dummy error, but not as a contract-costing error. Especially at IMP scoring it may be good to be aware of this. An alternative might be to weight the errors by IMP cost (so losing an overtrick costs only 1, but losing the contract costs 10 or whatever for a game).

(3) One can avoid contract-costing errors by constantly underbidding, so every contract rates to make +1 or +2.

(4) On defense, there will be impact from partner's signals/failure to signal to the degree that some mistakes will not be a defender's "fault."

However, despite the problems, this seems to me like a better way to measure skill than trying to look simply at end results.

If bridge is an art of partnership, it's just a big nonsense to rate individual's game
(unless in individual tournaments). If you rate partnership strength, it's extremely simple, about the same as chess.
0

#35 User is offline   BillHiggin 

  • PipPipPipPip
  • Group: Full Members
  • Posts: 499
  • Joined: 2007-February-03

Posted 2007-July-06, 11:09

junyi_zhu, on Jul 6 2007, 10:11 AM, said:

If you rate partnership strength, it's extremely simple, about the same as chess.

Rating partnerships in bridge is much more complicated and problematic than rating chess. Bridge results (the basis for a rating) are not based on your partnership's performance compared to the performance of the opposing partnership at the same table. Rather it is your partnership's performance compared to the performance of the set of other partnerships holding the same cards IN COMBINATION with the performance of the other partnership at your table compared to the set of other partnerships sitting their direction on the same deal. The concept of "set of other partnerships..." can be conveniently considered as the corresponding field (and is a single partnerships in pure team games), but still bridge rating involves 8 (virtual) players or 4 (virtual) partnerships rather than the simple 2 opponents of chess.

I have seen an attempt to adapt ELO ratings to bridge. The result was that the ratings did not behave similarily to chess ratings and more importantly the perception of those subject to the ratings was not positive. Chess ratings are transparent (you know how they are calculated and can verify the calculations) and most importantly they are percieved as fair and accurate. The perception issue is the toughest nut to crack. If those being rated have any perceptions that the rating is not accurate or overly subject to manipulation then the ratings become a source of trouble and complaints.
You must know the rules well - so that you may break them wisely!
0

#36 User is offline   bassaidai 

  • PipPip
  • Group: Members
  • Posts: 17
  • Joined: 2006-November-12

Posted 2007-July-06, 12:13

junyi_zhu, on Jul 6 2007, 04:03 PM, said:

GIB is generally weaker than intermediate players

I don't think so: how many boards did you play against/with GIB ?

How much thinking time did you allow GIB ?
Sapere aude !

I. Kant
0

#37 User is offline   mikeh 

  • PipPipPipPipPipPipPipPipPipPip
  • Group: Advanced Members
  • Posts: 13,862
  • Joined: 2005-June-15
  • Gender:Male
  • Location:Canada
  • Interests:Bridge, golf, wine (red), cooking, reading eclectically but insatiably, travelling, making bad posts.

Posted 2007-July-06, 12:57

Any attempt to rate players card play (defence or declarer) by a double-dummy analysis is not only silly, but completely misses the point.

While true experts will sometimes adopt the double-dummy line, that will actually constitute an error some of the time. The beauty of the game lies, in part, in the fact that the best single-dummy line doesn't always work.

This is obvious: holding AQJxx opposite 109xx, with no clues from the bidding, the correct single-dummy line is to finesse: but some of the time the correct double-dummy line is to play the A, dropping the stiff K offside.

No expert would do this under anything approaching normal circumstances and so taking a hook losing to a stiff K will be recorded as a double-dummy error.

Yet, if a ruff threatened, and we could afford to lose a trump trick (that suit as trump_ but not a trump trick and a ruff, we may well play the A rather than finesse. When the Kx(x)(x) is onside, we have committed a double-dummy error, altho, single-dummy, we made the right play.

Furthermore, at the table, there are real-life inferences available to an expert declarer or defender... inferences from bids or passes, based on one's assessment of the level of aggression of the opps, inferences based on tiny or marked breaks in tempo, etc.

I am not at all surprised that the double dummy analysis has .4 errors per board for WC... but I am willing to bet it is down to about .1 single dummy. At the same time, the few times I have kibb'd friends who are indifferent players or looked at samples of boards played many times on BBO, suggests to me that the average player makes more than 1 mistake per board... but that many of these mistakes don't cost, and would therefore not be caught by a computer analysis. A common type of error is the order in which we play the suits or the cards within a suit.. timing issues. And many of those end up not costing because the correct play caters to low-probability events. And so on. Not to mention the resemblance of bad bridge to watching tennis, as the errors go flying back and forth, often cancelling each other out B)
'one of the great markers of the advance of human kindness is the howls you will hear from the Men of God' Johann Hari
0

#38 User is offline   nige1 

  • 5-level belongs to me
  • PipPipPipPipPipPipPipPipPip
  • Group: Advanced Members
  • Posts: 9,128
  • Joined: 2004-August-30
  • Gender:Male
  • Location:Glasgow Scotland
  • Interests:Poems Computers

Posted 2007-July-06, 15:26

junyi_zhu, on Jul 6 2007, 10:03 AM, said:

World Class and GIB make 0.4 errors per hand in card plays? That's just an insult of world class players, hahaha.

I reckon that if you make less than 2 mistakes per board (bidding and play) then you are world-class :) :) :) You may even be a world-champion B) B) B)

I don't mean double-dummy errors or even sophisticated errors -- I mean ordinary face-to-face practical errors that most players would accept if simply explained. -- like daisy-picking in the bidding -- or failing to make a play that could help prevent partner misdefending.

At Bridge, few errors cost. Some actually gain because of a lucky lie of the cards or compensating opponent error.

For a season, our team conducted a detailed post-mortem after every match. Collectively we never chucked less than 100 imps even in 24 and 32 board matches :) Some of these matches, we won, in national competition. On a few occasions, opponents resigned with boards to play. In 1 or 2 matches, compensating errors meant that the score-card deceptively showed us conceding less than 2 imps per board. Although, of course, we should have lost much more :( :( :(
0

#39 User is offline   junyi_zhu 

  • PipPipPipPipPip
  • Group: Full Members
  • Posts: 536
  • Joined: 2003-May-28
  • Location:Saltlake City

Posted 2007-July-07, 18:18

bassaidai, on Jul 6 2007, 06:13 PM, said:

junyi_zhu, on Jul 6 2007, 04:03 PM, said:

GIB is generally weaker than intermediate players

I don't think so: how many boards did you play against/with GIB ?

How much thinking time did you allow GIB ?

A team of four with about 500 ACBL master points would beat a team of four gibs without much difficulties in a 64 board match, IMO, if both play sayc, under current BBO's slowest setting of gib. GIB has way more bugs than you can ever imagine. And it's certainly a wrong judgement of Zia to give up his famous bet when gib was invented.
0

#40 User is offline   junyi_zhu 

  • PipPipPipPipPip
  • Group: Full Members
  • Posts: 536
  • Joined: 2003-May-28
  • Location:Saltlake City

Posted 2007-July-07, 18:24

nige1, on Jul 6 2007, 09:26 PM, said:

junyi_zhu, on Jul 6 2007, 10:03 AM, said:

World Class and GIB make 0.4 errors per hand in card plays? That's just an insult of world class players, hahaha.

I reckon that if you make less than 2 mistakes per board (bidding and play) then you are world-class :) :) :) You may even be a world-champion B) B) B)

I don't mean double-dummy errors or even sophisticated errors -- I mean ordinary face-to-face practical errors that most players would accept if simply explained. -- like daisy-picking in the bidding -- or failing to make a play that could help prevent partner misdefending.

At Bridge, few errors cost. Some actually gain because of a lucky lie of the cards or compensating opponent error.

For a season, our team conducted a detailed post-mortem after every match. Collectively we never chucked less than 100 imps even in 24 and 32 board matches B) Some of these matches, we won, in national competition. On a few occasions, opponents resigned with boards to play. In 1 or 2 matches, compensating errors meant that the score-card deceptively showed us conceding less than 2 imps per board. Although, of course, we should have lost much more :) :( :(

I have watched and commented enough top level matches to claim that it's not even close to 0.4 errors per board for world class players at top level competitions. Most hands in bridge are routine and there exists a huge gap between intermediate and world class players. Intermediate players blow at least one trick per hand(some may not cost because errors can cancel out). For world class players, I agree with Mikeh, it's less than 0.1 erros per board.
0

  • 3 Pages +
  • 1
  • 2
  • 3
  • You cannot start a new topic
  • You cannot reply to this topic

1 User(s) are reading this topic
0 members, 1 guests, 0 anonymous users