Rating systems (again?)
#41
Posted 2007-July-07, 22:51
#42
Posted 2007-July-08, 00:08
(1) Measuring single dummy is hard. Measuring double dummy is relatively much easier.
(2) Some players are known for table feel, or for reading the opponents leads or bidding. This enables them to find "anti-percentage" lines that work. If they can really do this, that should be rewarded by the rating system rather than penalized because "they didn' find the best single-dummy line."
(3) It's nice that you can actually rate different aspects of people's play (declarer play, opening lead, defense, even distinguish between trump contracts and notrump or between partials and slams).
There is an issue with safety plays, since taking a safety play is usually a "double dummy mistake" but actually maximizes the expected score. Simple way to fix this is to count mistakes based on the change in result, so that a safety play "mistake" is just lose one but a drop the contract on the floor "mistake" is lose a lot more.
a.k.a. Appeal Without Merit
#43
Posted 2007-July-08, 03:25
It wouldn't surprise me too much if GIB is close to WC level. Jack is on par with the Dutch internationals against whom it has played team matches, and my personal impression is that Jack and GIB are at similar level. We have had this disucssion in other threads as well. I think people tend to under-estimate GIB.
Anyway, I don't see how one can compute error rates in an automatic way. You never know why a good player chooses an anti-percentage line - maybe it was actually a percentage line given the (misleading?) information he had about opps carding methods, BITs etc. Maybe he was trying a deceptive line. Maybe he was behaving ethically, playing contrary to what was suggested by UI he recieved. Maybe he was speculating about what might happen at the other table.
#44
Posted 2007-July-08, 03:41
This way I could use vugraph files and files from myhand archive of well known WC players.
To get a valid score more than 400 boards are needed (player will be dummy in about 100 of them).
Most GIB errors are lead errors or happen during the first tricks.
GIB errors that happen later are usually percentage plays that unfortunately don't work, while a simpler line would not have failed.
#45
Posted 2007-July-08, 12:11
helene_t, on Jul 8 2007, 09:25 AM, said:
It wouldn't surprise me too much if GIB is close to WC level. Jack is on par with the Dutch internationals against whom it has played team matches, and my personal impression is that Jack and GIB are at similar level. We have had this disucssion in other threads as well. I think people tend to under-estimate GIB.
Anyway, I don't see how one can compute error rates in an automatic way. You never know why a good player chooses an anti-percentage line - maybe it was actually a percentage line given the (misleading?) information he had about opps carding methods, BITs etc. Maybe he was trying a deceptive line. Maybe he was behaving ethically, playing contrary to what was suggested by UI he recieved. Maybe he was speculating about what might happen at the other table.
You were talking about Pairs and top dutch players might not be familiar with jack. For a long team match, any computer programs simply have no chance against a top flight team of humanbeing, cause those programs all have numerous bugs and humanbeing can explore their weakness and take the full advantage of their bugs. The declarer play of gib is slightly better than intermediate human players (they don't usually understand safe plays, cause those are rare events and hard to produce in a limit number of random deals; they don't understand opp's bidding, cause it's extremely hard to set those right constraints, which requires a big advance in AI), the bidding of gib is way worse than intermediate human players,, that means gibs tend to make very costly mistakes in bidding and we all know how important bidding is in bridge. Also the truth is not that most people underrate computer, the situation is that the programmers often overate the programs IMO. I am only trying to be objective after played thousands of hands with gib and few in this forum may have played the same number of hands as I have done.
#46
Posted 2007-July-08, 12:11
awm, on Jul 5 2007, 06:17 PM, said:
(1) A play is a double-dummy error if it reduces the double-dummy trick total for the person making the play. It's easy to notice these when kibitzing using GIB.
(2) A play is a double-dummy contract-costing error if it changes the result of the hand from making to down (for declarer) or vice versa (for defense).
In principle we could compute these error rates. This has a number of nice features; in particular the error rates for declarer are mostly independent of partner. There are obviously a few things to watch out for, in particular:
hotShot, on Jul 5 2007, 11:30 AM, said:
It is less significant as one might think. WC's and gib produce error rates of about 0.4 errors per board. Intermediates make about 0.8 and beginners reach 0.9+.
I had a few pickup partner with rates larger than 1, but i did not play enough boards with them to make the results mean something in a statistic.
hotShot, on Jul 8 2007, 04:41 AM, said:
This way I could use vugraph files and files from myhand archive of well known WC players.
To get a valid score more than 400 boards are needed (player will be dummy in about 100 of them).
Most GIB errors are lead errors or happen during the first tricks.
GIB errors that happen later are usually percentage plays that unfortunately don't work, while a simpler line would not have failed.
Fascinating stuff, hotShot! Thank you! I am amazed. I find it hard to believe that the error rate is so low! Please tell us how you count a player's errors on one board? In particular, what is the range and variance for different categories of player?
I recckon that Double-dummy analysis is rather cruel even if it can't catch all the mistakes picked up by Single-dummy analysis. For example...
- Signaling errors.
- Failures to help partner.
- Failures to give an opponent a guess.
Thus failure to drop a singleton king offside is a "mistake" at double-dummy.
Also, for example, Scoring considerations affect correct single-dummy play. Often declarer makes the "mistake" of adopting a safer single-dummy line when a riskier double-dummy line would have garnered more over-tricks. Similarly, a defender, at teams, often sacrifices potential extra under-tricks to achieve his prime target of defeating the contract.
In my experience, players often make several "mistakes" on a single board. For instance, a defender makes a trick-costing "mistake" -- but declarer promptly makes a compensating "mistake". Sometimes this happens again and again on the same board. Occasionally, in spite of this comedy of errors, the end-result seems quite normal!
#47
Posted 2007-July-09, 01:06
For GIB: My wife started to play just some month ago and she decided that she would be a plague to any human opponent. So we played the gibs. I think the gibs are far better then the intermedeate players around here. (We had
And they are better then the normal club players.
But they are lightyears away from Worldclass.
I have big respect for dutch nationalist, so I guess that Jack ran on a much better engine and with other settings then the gib that is working here.
As far as I know, people who own GIB found out that he is playing better in the normal buyable programm then he is playing here.
Roland
Sanity Check: Failure (Fluffy)
More system is not the answer...
#48
Posted 2007-July-19, 06:47
1. When calculating errors, it is ok to say that failing to drop a stiff king missing 3 in the suit is an error, 'everyone' makes that error, and we are comparing rates, so it is not a big deal.
2. A World Champion once told me if you make only 2 mistakes in a session, you have played very well. So, I doubt the WC players make less than 0.1 mistakes in a session.
3. Everyone so far is missing a big point. Bridge is at least 1/2 bidding. Any rating system that tries to count errors needs to look at bidding as well. This would be an almost impossible task. For example, we would call it an error to bid a slam on 2 finesses. But what if they were through an opening bidder? Or into a preemptors hand? What about auctions that go, say 3s-6s, and the leader has to guess the suit? Impossible to evaluate. Is it an error to preempt, catch partner with the death 4450 with a void in your suit? Of course not.
4. There are plenty of mistakes that do not appear as mistakes, as many have pointed out. Failing to cater to an offside stiff queen is an easy example. Sloppy signalling is another. In the bidding, making a bid that partner doesn't understand is yet another.
Danny
#49
Posted 2007-July-19, 07:02
GIB was the first of a new generation of computer bridge programs and at the time miles ahead of the rest. The Jack team realized that the biggest gain was still in the auction (although I'm sure Jack is also better in the card play than GIB, simply because it was based on it and some improvements have been made since then).
From my personal experience: Playing the money bridge tourneys, I end up positive long term playing total points against 3 GIBs . This must mean my total playing strength is higher than GIBs(at least the version that is implemented here). On the other hand I suspect I would lose long term playing total points with 3 Jacks.
#50
Posted 2007-July-19, 07:53
hotShot, on Jul 5 2007, 11:30 AM, said:
Wow, I'd estimate I have 2 errors per board in signalling alone. Telling partner what he doesn't need to know (which in theory might tell declarer what she needs to know) and not telling partner what to save for the endgame are the big two. Of course, most of the time these don't make any difference: partner can figure out what to save on his own, and even if declarer's paying attention the information isn't of any use to her.
How do you measure errors?
#51
Posted 2007-July-19, 07:55
Quote
GIB was the first of a new generation of computer bridge programs and at the time miles ahead of the rest. The Jack team realized that the biggest gain was still in the auction (although I'm sure Jack is also better in the card play than GIB, simply because it was based on it and some improvements have been made since then).
It's interesting you mention that... I wound up buying BridgeBaron a handful of months ago when I started to get serious about learning.
At least part of the motivation to pick that product was based on looking at the results of one of Richard Pavilcek's bidding polls (like this, for example).
#52
Posted 2007-July-19, 08:03
Gerben42, on Jul 19 2007, 03:02 PM, said:
The Jack team claims that they have improved the defensive card play substantially. I haven't played that much with version 4 but I agree with their observation that the defensive play was prone to improvement in earlier versions.
#53
Posted 2007-July-19, 08:21
jtfanclub, on Jul 19 2007, 03:53 PM, said:
I recalculate the number of tricks that NS can make after each card played.
If the number after the play is smaller NS made an "error", if it's greater EW must have made one.
The blame always goes to the one who played the card, although because of signals, lead directing doubles and things like that, the real responsibility must be given to his partner.
#54
Posted 2007-July-19, 08:36
hotShot, on Jul 19 2007, 09:21 AM, said:
jtfanclub, on Jul 19 2007, 03:53 PM, said:
I recalculate the number of tricks that NS can make after each card played.
If the number after the play is smaller NS made an "error", if it's greater EW must have made one.
The blame always goes to the one who played the card, although because of signals, lead directing doubles and things like that, the real responsibility must be given to his partner.
Ah, well then. Luckily, most signalling errors don't do any harm, so they won't factor into your calculation of errors.
Heck, I once played a game where we agreed on up-side down attitude on signals and discards and then partner forgot for the entire game and signalled normally (and took my signals as normal). It made a difference once in 12 boards.
#55
Posted 2007-July-19, 09:05
Quote
Although GIB is not world class, it was revolutionary and solved some problems in computer bridge that allows in principle world class computer programs, so in a way Zia was correct recognizing that.
#56
Posted 2007-July-19, 10:05
You can't really rate an individual's bidding acumen by looking at hands. If you put two "good bidders" together but they're accustomed to different agreements and style, they're likely to have a lot of problems and reach a lot of poor contracts. If you put two "mediocre bidders" together but they're a well-established partnership with lots of agreements (and quite possibly their notes in front of them if playing online), they will do quite well.
If you want to rate bidding, I think you have to do it by partnership.
a.k.a. Appeal Without Merit
#57
Posted 2007-July-19, 14:10
#58
Posted 2007-July-20, 04:50
hotShot, on Jul 19 2007, 09:21 AM, said:
jtfanclub, on Jul 19 2007, 03:53 PM, said:
I recalculate the number of tricks that NS can make after each card played. If the number after the play is smaller NS made an "error", if it's greater EW must have made one.
Thank you hotshot.
I am amazed that GIB and "World-class" players average less that one "double-dummy" play-error in every two boards.
I do accept your figures but I'm surprised. I would have thought that, for example, an expert declarer could often finesse the wrong way; or finesse when the drop would have worked. At teams, he would also play to make his contracts when a successful risky play for overtricks was available.
I wold expect defenders to make even more "mistakes", especially on the opening lead.
More questions...
- Does the average include when a player is dummy?
- How do defender and declarer "errors" compare?
- In particular, what is the range and variance of such "errors". For example, do you find many boards with more than 5 errors
#59
Posted 2007-July-22, 14:03
Well the software keep extra records for errors in declarer play, defense and leads, but I did not investigate much. I don't have enough recorded boards for a full/serious statistic analysis. If I have 400 recorded deals with a player, he's dummy in 100, so they are lost. He's declarer in about 100 and defender in about 200, which is not much for serious analysis.
#60
Posted 2007-July-22, 23:09
hotShot, on Jul 22 2007, 03:03 PM, said:
Well the software keep extra records for errors in declarer play, defense and leads, but I did not investigate much. I don't have enough recorded boards for a full/serious statistic analysis. If I have 400 recorded deals with a player, he's dummy in 100, so they are lost. He's declarer in about 100 and defender in about 200, which is not much for serious analysis.
Oh Well. Thanks anyway. You are lucky with your pick-up partners hotshot

Help