\

Orca-Bench: How Ready Are Language Model Agents for Oncall?

18 points - today at 6:32 PM

Source
  • 4di

    today at 7:25 PM

    looks like the public bench link in the paper was taken down. https://hub.harborframework.com/datasets/orca-bench/ORCA-ben...

    This doesn't work anymore. Is there a newer link?

    • dash2

      today at 7:01 PM

      Seems like there's a big attack-defence asymmetry at present: models are great at exploiting systems and poor at fixing them.

        • aleksiy123

          today at 8:12 PM

          Attackers advantage in the iterative fast feedback loop?

          It’s harder to have a loop to ensure you are defending all possible attacks?

          I guess the loop is you need to attack yourself and fix. But attackers only need a single opening.

          Finding all possible attacks and patching them against yourself is inherently more expensive?

          • EGreg

            today at 7:22 PM

            That is why I built https://safebots.ai/safebox.html

            Your strategy can’t be patch AFTER an intrusion. Only to build a hardened environment from scratch and be ready in advance.