• jj4211@lemmy.world
    link
    fedilink
    arrow-up
    5
    ·
    13 hours ago

    I have both used agents and have been the “victim” of heavy agent users.

    If you can provide an utterly well bounded problem with perfectly verifiable criteria for it to retry against until the tests pass and the tests are full and valid for the use case, it can work. Similarly, if your code’s reality is forgiving and flexible, and a ‘close enough’ result is good enough to get the human software user in the right ball park, you might be able to extract decent behavior.

    However, if there is a means by which the answer can ‘look correct’ yet be wrong in the real world and the real world scenarios require accuracy and precision, there’s huge gaps.

    Especially if the agents control the coding and the test case generation, seen plenty of times where it talked itself out of a test case that was failing when the test case was in fact correctly showing a flaw.

    • iLigator@lemmy.zip
      link
      fedilink
      arrow-up
      1
      ·
      12 hours ago

      That’s actually my biggest issue with it, it often tries to get out of scope and touch something it shouldn’t. But thats a low bar for an entire profession to exist. I wish llms never happened, ruined our field and I’m upset about that. I honestly cannot imagine a CS field it wont disrupt.