Hacker Times
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
ltononro
31 days ago
|
parent
|
context
|
favorite
| on:
Senior SWE-Bench: open-source benchmark that asses...
Using LLMs to judge this could be very hackable, shouldnt it be? Is that the best practice we expect? Would be very interested in see some ablations about failure modes where the LLMs tried to hack it somehow. Or failures from the llmaaj
Consider applying for YC's Fall 2026 batch!
Applications
are open till July 27.
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: