\

When LLM judges agree, should we believe them?

32 points - today at 4:29 PM

Source
  • ex1fm3ta

    today at 7:03 PM

    I kinda find it funny when I use the advisor on claude code and it agrees with the ideas that the previous model did.

    For info: the advisor(s) available are higher end models. For example: you use sonnet, the available advisors are opus and fable. If you use Haiku, the advisor are sonnet, opus and fable.

    • qarl

      today at 6:19 PM

      While this is absolutely true - I'd hesitate to discount using similar agents for checking each other. Two agents will almost never hallucinate in the same way, regardless of their weights - and by having a second one (with a different context) check almost entirely eliminates the problem.

        • emodendroket

          today at 7:08 PM

          It depends what we're judging, doesn't it? If it's "is the formatting in this document compliant with our standards?" I think it's reasonable. If it's like, life-altering if it's wrong I'm less sanguine.

      • bryzaguy

        today at 6:07 PM

        They would all agree raspberry has two Rs

          • Joel_Mckay

            today at 6:59 PM

            But still refuse to answer "How many strings does a bass play with in water?" , perhaps the chat monitors in the third world data entry centers will manually patch the nonsense for a more rational answer someday. lol =3

        • VaradD09

          today at 6:23 PM

          I believe it depends on the LLM itself. Like what model as each model has diff weights and diff data trained onn

            • dgellow

              today at 6:37 PM

              I would recommend to read the article, it’s actually more nuanced than the title

          • Tsarp

            today at 5:11 PM

            Kinda weird to generalize "LLM". Every lab, every model is different. Has its own biases, reward functions etc.

              • Centigonal

                today at 5:42 PM

                [dead]

            • Founderarcstone

              today at 6:10 PM

              Great point this will be interesting how this develops.

              • troupo

                today at 6:04 PM

                Without reading the article (doesn't matter if it's pro or contra): no, of course not.

                It shouldn't even be a debatable question.

                  • dgellow

                    today at 6:36 PM

                    I think you should have read the article first, at minimum the subheader

                    > Discounting the opinions of LLM judges with highly correlated outputs ensures that panels of judges reflect a true diversity of perspectives.

                • novaapi

                  today at 5:35 PM

                  [dead]

                  • asamoahf

                    today at 4:59 PM

                    The unsupervised framing is the part I'd push on. If true labels are latent and you infer them jointly with judge parameters, then a blind spot every judge shares isn't a correlated error the model can down-weight. It's indistinguishable from the ground truth, and the likelihood has no reason to prefer the correct answer over the consensus one.

                    So this fixes dependence between judges and leaves dependence between all the judges and the truth untouched, which is the failure people are actually worried about when they say eight models agreed. You still want a small human-labelled anchor set to break it. The number I'd find interesting is how much smaller that anchor set gets once you model the dependence, since that's the real saving.

                    Same shape as offline policy evaluation. Correlated logging errors survive any amount of re-weighting, and one real experiment would be probably what pins them.

                      • Forgeties79

                        today at 5:56 PM

                        the LLM-speak is unbearable

                          • vips7L

                            today at 6:46 PM

                            It’s awful. A ton of words to say absolutely nothing.