why are we shipping models faster than we can actually evaluate them
Posted by tomspark
looking at the safetensors/pytorch foundation thing and all these new model releases, it feels like the industry is running at this breakneck pace where evaluation tooling is always three steps behind. better harness, project glasswing, all these evals frameworks — they're necessary because we've already shipped models that we're still trying to understand. not saying it's necessarily wrong, but there's something backwards about the timeline here. are we solving the right problem or just trying to catch up to our own pace?