BBO Discussion Forums: Rating systems (again?) - BBO Discussion Forums

Jump to content

  • 3 Pages +
  • 1
  • 2
  • 3
  • You cannot start a new topic
  • You cannot reply to this topic

Rating systems (again?)

#41 User is offline   DKJ 

  • Pip
  • Group: Members
  • Posts: 1
  • Joined: 2007-July-07

Posted 2007-July-07, 22:51

It is so very obvious that the current practice of the players descibing their own level of bridge skill is seriously detrimental to playing on BBO; there is simply on dearth of deluded ones and 'jokers'!. A more objective system must be introduced soonest possible, sesily done by BBO point system, or a rating system by partners in actual play...or a combination of both.
0

#42 User is offline   awm 

  • PipPipPipPipPipPipPipPipPip
  • Group: Advanced Members
  • Posts: 8,707
  • Joined: 2005-February-09
  • Gender:Male
  • Location:Zurich, Switzerland

Posted 2007-July-08, 00:08

There is obviously a difference between single dummy errors and double dummy errors. I agree that top flight players make quite few single dummy errors (other than possibly on the opening lead). However, this doesn't necessarily invalidate the idea of measuring double dummy errors. If you routinely take the best single dummy line, sometimes you will do the "wrong thing" double dummy, but more often than not you will do the "right thing." In the long run, this will tend to average out and people who are taking the best single dummy line will end up with fewer (but nonzero) "errors per board." Obviously no one (except a cheater) will end with zero "errors per board." A number of nice aspects to this:

(1) Measuring single dummy is hard. Measuring double dummy is relatively much easier.
(2) Some players are known for table feel, or for reading the opponents leads or bidding. This enables them to find "anti-percentage" lines that work. If they can really do this, that should be rewarded by the rating system rather than penalized because "they didn' find the best single-dummy line."
(3) It's nice that you can actually rate different aspects of people's play (declarer play, opening lead, defense, even distinguish between trump contracts and notrump or between partials and slams).

There is an issue with safety plays, since taking a safety play is usually a "double dummy mistake" but actually maximizes the expected score. Simple way to fix this is to count mistakes based on the change in result, so that a safety play "mistake" is just lose one but a drop the contract on the floor "mistake" is lose a lot more.
Adam W. Meyerson
a.k.a. Appeal Without Merit
0

#43 User is offline   helene_t 

  • The Abbess
  • PipPipPipPipPipPipPipPipPipPipPip
  • Group: Advanced Members
  • Posts: 17,401
  • Joined: 2004-April-22
  • Gender:Female
  • Location:Odense, Denmark
  • Interests:History, languages

Posted 2007-July-08, 03:25

What kind of World Class players make those 0.4 errors? The self-rated ones or the real ones?

It wouldn't surprise me too much if GIB is close to WC level. Jack is on par with the Dutch internationals against whom it has played team matches, and my personal impression is that Jack and GIB are at similar level. We have had this disucssion in other threads as well. I think people tend to under-estimate GIB.

Anyway, I don't see how one can compute error rates in an automatic way. You never know why a good player chooses an anti-percentage line - maybe it was actually a percentage line given the (misleading?) information he had about opps carding methods, BITs etc. Maybe he was trying a deceptive line. Maybe he was behaving ethically, playing contrary to what was suggested by UI he recieved. Maybe he was speculating about what might happen at the other table.
The world would be such a happy place, if only everyone played Acol :) --- TramTicket
0

#44 User is offline   hotShot 

  • Axxx Axx Axx Axx
  • PipPipPipPipPipPipPip
  • Group: Advanced Members
  • Posts: 2,976
  • Joined: 2003-August-31
  • Gender:Male

Posted 2007-July-08, 03:41

My program scans all lin-files in a given directory, to perform the analysis.
This way I could use vugraph files and files from myhand archive of well known WC players.
To get a valid score more than 400 boards are needed (player will be dummy in about 100 of them).

Most GIB errors are lead errors or happen during the first tricks.
GIB errors that happen later are usually percentage plays that unfortunately don't work, while a simpler line would not have failed.
0

#45 User is offline   junyi_zhu 

  • PipPipPipPipPip
  • Group: Full Members
  • Posts: 536
  • Joined: 2003-May-28
  • Location:Saltlake City

Posted 2007-July-08, 12:11

helene_t, on Jul 8 2007, 09:25 AM, said:

What kind of World Class players make those 0.4 errors? The self-rated ones or the real ones?

It wouldn't surprise me too much if GIB is close to WC level. Jack is on par with the Dutch internationals against whom it has played team matches, and my personal impression is that Jack and GIB are at similar level. We have had this disucssion in other threads as well. I think people tend to under-estimate GIB.

Anyway, I don't see how one can compute error rates in an automatic way. You never know why a good player chooses an anti-percentage line - maybe it was actually a percentage line given the (misleading?) information he had about opps carding methods, BITs etc. Maybe he was trying a deceptive line. Maybe he was behaving ethically, playing contrary to what was suggested by UI he recieved. Maybe he was speculating about what might happen at the other table.

You were talking about Pairs and top dutch players might not be familiar with jack. For a long team match, any computer programs simply have no chance against a top flight team of humanbeing, cause those programs all have numerous bugs and humanbeing can explore their weakness and take the full advantage of their bugs. The declarer play of gib is slightly better than intermediate human players (they don't usually understand safe plays, cause those are rare events and hard to produce in a limit number of random deals; they don't understand opp's bidding, cause it's extremely hard to set those right constraints, which requires a big advance in AI), the bidding of gib is way worse than intermediate human players,, that means gibs tend to make very costly mistakes in bidding and we all know how important bidding is in bridge. Also the truth is not that most people underrate computer, the situation is that the programmers often overate the programs IMO. I am only trying to be objective after played thousands of hands with gib and few in this forum may have played the same number of hands as I have done.
0

#46 User is offline   nige1 

  • 5-level belongs to me
  • PipPipPipPipPipPipPipPipPip
  • Group: Advanced Members
  • Posts: 9,128
  • Joined: 2004-August-30
  • Gender:Male
  • Location:Glasgow Scotland
  • Interests:Poems Computers

Posted 2007-July-08, 12:11

awm, on Jul 5 2007, 06:17 PM, said:

However, on BBO we really have more information than just the net results of contracts. We could analyze each card played. In particular, I like the idea of computing double-dummy error rate. We could define:

(1) A play is a double-dummy error if it reduces the double-dummy trick total for the person making the play. It's easy to notice these when kibitzing using GIB.

(2) A play is a double-dummy contract-costing error if it changes the result of the hand from making to down (for declarer) or vice versa (for defense).

In principle we could compute these error rates. This has a number of nice features; in particular the error rates for declarer are mostly independent of partner. There are obviously a few things to watch out for, in particular:

hotShot, on Jul 5 2007, 11:30 AM, said:

I have done that and posted some results of that a while ago. But the system is limited to play and ignores bidding completely.
It is less significant as one might think. WC's and gib produce error rates of about 0.4 errors per board. Intermediates make about 0.8 and beginners reach 0.9+.
I had a few pickup partner with rates larger than 1, but i did not play enough boards with them to make the results mean something in a statistic.

hotShot, on Jul 8 2007, 04:41 AM, said:

My program scans all lin-files in a given directory, to perform the analysis.
This way I could use vugraph files and files from myhand archive of well known WC players.
To get a valid score more than 400 boards are needed (player will be dummy in about 100 of them).

Most GIB errors are lead errors or happen during the first tricks.
GIB errors that happen later are usually percentage plays that unfortunately don't work, while a simpler line would not have failed.


Fascinating stuff, hotShot! Thank you! I am amazed. I find it hard to believe that the error rate is so low! Please tell us how you count a player's errors on one board? In particular, what is the range and variance for different categories of player?

I recckon that Double-dummy analysis is rather cruel even if it can't catch all the mistakes picked up by Single-dummy analysis. For example...
  • Signaling errors.
  • Failures to help partner.
  • Failures to give an opponent a guess.
Double-dummy criteria may be more objective than Single-dummy but I would expect them to high-light a similar number of errors, although, obviously, a "mistake" at double dummy is sometimes the correct play at single-dummy.

Thus failure to drop a singleton king offside is a "mistake" at double-dummy.

Also, for example, Scoring considerations affect correct single-dummy play. Often declarer makes the "mistake" of adopting a safer single-dummy line when a riskier double-dummy line would have garnered more over-tricks. Similarly, a defender, at teams, often sacrifices potential extra under-tricks to achieve his prime target of defeating the contract.

In my experience, players often make several "mistakes" on a single board. For instance, a defender makes a trick-costing "mistake" -- but declarer promptly makes a compensating "mistake". Sometimes this happens again and again on the same board. Occasionally, in spite of this comedy of errors, the end-result seems quite normal!
0

#47 User is offline   Codo 

  • PipPipPipPipPipPipPipPip
  • Group: Advanced Members
  • Posts: 6,373
  • Joined: 2003-March-15
  • Location:Hamburg, Germany
  • Interests:games and sports, esp. bridge,chess and (beach-)volleyball

Posted 2007-July-09, 01:06

For the rating stuff: We had these discussion before, I have nothing to add.

For GIB: My wife started to play just some month ago and she decided that she would be a plague to any human opponent. So we played the gibs. I think the gibs are far better then the intermedeate players around here. (We had
And they are better then the normal club players.
But they are lightyears away from Worldclass.
I have big respect for dutch nationalist, so I guess that Jack ran on a much better engine and with other settings then the gib that is working here.
As far as I know, people who own GIB found out that he is playing better in the normal buyable programm then he is playing here.
Kind Regards

Roland


Sanity Check: Failure (Fluffy)
More system is not the answer...
0

#48 User is offline   3for3 

  • PipPipPip
  • Group: Full Members
  • Posts: 93
  • Joined: 2004-August-26

Posted 2007-July-19, 06:47

Some thoughts about this thread.

1. When calculating errors, it is ok to say that failing to drop a stiff king missing 3 in the suit is an error, 'everyone' makes that error, and we are comparing rates, so it is not a big deal.

2. A World Champion once told me if you make only 2 mistakes in a session, you have played very well. So, I doubt the WC players make less than 0.1 mistakes in a session.

3. Everyone so far is missing a big point. Bridge is at least 1/2 bidding. Any rating system that tries to count errors needs to look at bidding as well. This would be an almost impossible task. For example, we would call it an error to bid a slam on 2 finesses. But what if they were through an opening bidder? Or into a preemptors hand? What about auctions that go, say 3s-6s, and the leader has to guess the suit? Impossible to evaluate. Is it an error to preempt, catch partner with the death 4450 with a void in your suit? Of course not.

4. There are plenty of mistakes that do not appear as mistakes, as many have pointed out. Failing to cater to an offside stiff queen is an easy example. Sloppy signalling is another. In the bidding, making a bid that partner doesn't understand is yet another.


Danny
0

#49 User is offline   Gerben42 

  • PipPipPipPipPipPipPipPip
  • Group: Advanced Members
  • Posts: 5,577
  • Joined: 2005-March-01
  • Gender:Male
  • Location:Erlangen, Germany
  • Interests:Astronomy, Mathematics
    Nuclear power

Posted 2007-July-19, 07:02

From what I've seen, Jack would massacre GIB in the bidding. Computer card play theory however has not improved much since GIB other than increase in computer strength.

GIB was the first of a new generation of computer bridge programs and at the time miles ahead of the rest. The Jack team realized that the biggest gain was still in the auction (although I'm sure Jack is also better in the card play than GIB, simply because it was based on it and some improvements have been made since then).

From my personal experience: Playing the money bridge tourneys, I end up positive long term playing total points against 3 GIBs . This must mean my total playing strength is higher than GIBs(at least the version that is implemented here). On the other hand I suspect I would lose long term playing total points with 3 Jacks.
Two wrongs don't make a right, but three lefts do!
My Bridge Systems Page

BC Kultcamp Rieneck
0

#50 User is offline   jtfanclub 

  • PipPipPipPipPipPipPip
  • Group: Advanced Members
  • Posts: 3,937
  • Joined: 2004-June-05

Posted 2007-July-19, 07:53

hotShot, on Jul 5 2007, 11:30 AM, said:

It is less significant as one might think. WC's and gib produce error rates of about 0.4 errors per board. Intermediates make about 0.8 and beginners reach 0.9+.

Wow, I'd estimate I have 2 errors per board in signalling alone. Telling partner what he doesn't need to know (which in theory might tell declarer what she needs to know) and not telling partner what to save for the endgame are the big two. Of course, most of the time these don't make any difference: partner can figure out what to save on his own, and even if declarer's paying attention the information isn't of any use to her.

How do you measure errors?
0

#51 User is offline   HeavyDluxe 

  • PipPipPipPip
  • Group: Full Members
  • Posts: 297
  • Joined: 2005-June-23
  • Gender:Male
  • Location:Windsor, VT

Posted 2007-July-19, 07:55

Quote

From what I've seen, Jack would massacre GIB in the bidding. Computer card play theory however has not improved much since GIB other than increase in computer strength.

GIB was the first of a new generation of computer bridge programs and at the time miles ahead of the rest. The Jack team realized that the biggest gain was still in the auction (although I'm sure Jack is also better in the card play than GIB, simply because it was based on it and some improvements have been made since then).


It's interesting you mention that... I wound up buying BridgeBaron a handful of months ago when I started to get serious about learning.

At least part of the motivation to pick that product was based on looking at the results of one of Richard Pavilcek's bidding polls (like this, for example).
0

#52 User is offline   helene_t 

  • The Abbess
  • PipPipPipPipPipPipPipPipPipPipPip
  • Group: Advanced Members
  • Posts: 17,401
  • Joined: 2004-April-22
  • Gender:Female
  • Location:Odense, Denmark
  • Interests:History, languages

Posted 2007-July-19, 08:03

Gerben42, on Jul 19 2007, 03:02 PM, said:

From what I've seen, Jack would massacre GIB in the bidding. Computer card play theory however has not improved much since GIB other than increase in computer strength.

The Jack team claims that they have improved the defensive card play substantially. I haven't played that much with version 4 but I agree with their observation that the defensive play was prone to improvement in earlier versions.
The world would be such a happy place, if only everyone played Acol :) --- TramTicket
0

#53 User is offline   hotShot 

  • Axxx Axx Axx Axx
  • PipPipPipPipPipPipPip
  • Group: Advanced Members
  • Posts: 2,976
  • Joined: 2003-August-31
  • Gender:Male

Posted 2007-July-19, 08:21

jtfanclub, on Jul 19 2007, 03:53 PM, said:

How do you measure errors?

I recalculate the number of tricks that NS can make after each card played.
If the number after the play is smaller NS made an "error", if it's greater EW must have made one.

The blame always goes to the one who played the card, although because of signals, lead directing doubles and things like that, the real responsibility must be given to his partner.
0

#54 User is offline   jtfanclub 

  • PipPipPipPipPipPipPip
  • Group: Advanced Members
  • Posts: 3,937
  • Joined: 2004-June-05

Posted 2007-July-19, 08:36

hotShot, on Jul 19 2007, 09:21 AM, said:

jtfanclub, on Jul 19 2007, 03:53 PM, said:

How do you measure errors?

I recalculate the number of tricks that NS can make after each card played.
If the number after the play is smaller NS made an "error", if it's greater EW must have made one.

The blame always goes to the one who played the card, although because of signals, lead directing doubles and things like that, the real responsibility must be given to his partner.

Ah, well then. Luckily, most signalling errors don't do any harm, so they won't factor into your calculation of errors.

Heck, I once played a game where we agreed on up-side down attitude on signals and discards and then partner forgot for the entire game and signalled normally (and took my signals as normal). It made a difference once in 12 boards.
0

#55 User is offline   Gerben42 

  • PipPipPipPipPipPipPipPip
  • Group: Advanced Members
  • Posts: 5,577
  • Joined: 2005-March-01
  • Gender:Male
  • Location:Erlangen, Germany
  • Interests:Astronomy, Mathematics
    Nuclear power

Posted 2007-July-19, 09:05

Quote

And it's certainly a wrong judgement of Zia to give up his famous bet when gib was invented.


Although GIB is not world class, it was revolutionary and solved some problems in computer bridge that allows in principle world class computer programs, so in a way Zia was correct recognizing that.
Two wrongs don't make a right, but three lefts do!
My Bridge Systems Page

BC Kultcamp Rieneck
0

#56 User is offline   awm 

  • PipPipPipPipPipPipPipPipPip
  • Group: Advanced Members
  • Posts: 8,707
  • Joined: 2005-February-09
  • Gender:Male
  • Location:Zurich, Switzerland

Posted 2007-July-19, 10:05

I'd make the point that there aren't really "great bidders." Bidding is so much a function of partnership and system. Take the best player in the world and force him or her to play a totally unfamiliar system and we'll see a lot of trouble in the bidding.

You can't really rate an individual's bidding acumen by looking at hands. If you put two "good bidders" together but they're accustomed to different agreements and style, they're likely to have a lot of problems and reach a lot of poor contracts. If you put two "mediocre bidders" together but they're a well-established partnership with lots of agreements (and quite possibly their notes in front of them if playing online), they will do quite well.

If you want to rate bidding, I think you have to do it by partnership.
Adam W. Meyerson
a.k.a. Appeal Without Merit
0

#57 User is offline   helene_t 

  • The Abbess
  • PipPipPipPipPipPipPipPipPipPipPip
  • Group: Advanced Members
  • Posts: 17,401
  • Joined: 2004-April-22
  • Gender:Female
  • Location:Odense, Denmark
  • Interests:History, languages

Posted 2007-July-19, 14:10

I don't think so. Take two World Class players with very different systems and style, say Sabine Auken and Huub Bertens, and they will just play some middle-of-the-road expert style which they both know reasonably well. Of course partnership harmony will not be the same as with their regular pd's, but it will still be quite good. I've seen some casual partnerships doing well in top competition. Last year one of the bigger Dutch tourneys was won by Bauke Muller and Jan Jansma. Marion MIchielsen has played successfully with just about everybody in the Dutch sub-top.
The world would be such a happy place, if only everyone played Acol :) --- TramTicket
0

#58 User is offline   nige1 

  • 5-level belongs to me
  • PipPipPipPipPipPipPipPipPip
  • Group: Advanced Members
  • Posts: 9,128
  • Joined: 2004-August-30
  • Gender:Male
  • Location:Glasgow Scotland
  • Interests:Poems Computers

Posted 2007-July-20, 04:50

hotShot, on Jul 19 2007, 09:21 AM, said:

jtfanclub, on Jul 19 2007, 03:53 PM, said:

How do you measure errors?

I recalculate the number of tricks that NS can make after each card played. If the number after the play is smaller NS made an "error", if it's greater EW must have made one.

Thank you hotshot.

I am amazed that GIB and "World-class" players average less that one "double-dummy" play-error in every two boards.

I do accept your figures but I'm surprised. I would have thought that, for example, an expert declarer could often finesse the wrong way; or finesse when the drop would have worked. At teams, he would also play to make his contracts when a successful risky play for overtricks was available.

I wold expect defenders to make even more "mistakes", especially on the opening lead.

More questions...
  • Does the average include when a player is dummy?
  • How do defender and declarer "errors" compare?
  • In particular, what is the range and variance of such "errors". For example, do you find many boards with more than 5 errors

0

#59 User is offline   hotShot 

  • Axxx Axx Axx Axx
  • PipPipPipPipPipPipPip
  • Group: Advanced Members
  • Posts: 2,976
  • Joined: 2003-August-31
  • Gender:Male

Posted 2007-July-22, 14:03

Board where a player is dummy are eliminated.
Well the software keep extra records for errors in declarer play, defense and leads, but I did not investigate much. I don't have enough recorded boards for a full/serious statistic analysis. If I have 400 recorded deals with a player, he's dummy in 100, so they are lost. He's declarer in about 100 and defender in about 200, which is not much for serious analysis.
0

#60 User is offline   nige1 

  • 5-level belongs to me
  • PipPipPipPipPipPipPipPipPip
  • Group: Advanced Members
  • Posts: 9,128
  • Joined: 2004-August-30
  • Gender:Male
  • Location:Glasgow Scotland
  • Interests:Poems Computers

Posted 2007-July-22, 23:09

hotShot, on Jul 22 2007, 03:03 PM, said:

Board where a player is dummy are eliminated.
Well the software keep extra records for errors in declarer play, defense and leads, but I did not investigate much. I don't have enough recorded boards for a full/serious statistic analysis. If I have 400 recorded deals with a player, he's dummy in 100, so they are lost. He's declarer in about 100 and defender in about 200, which is not much for serious analysis.

Oh Well. Thanks anyway. You are lucky with your pick-up partners hotshot :) Mine often make three or four trick costing errors on a single board :D Fortunately opponents often chuck the tricks right back :)
0

  • 3 Pages +
  • 1
  • 2
  • 3
  • You cannot start a new topic
  • You cannot reply to this topic

1 User(s) are reading this topic
0 members, 1 guests, 0 anonymous users