OpenAI פרסם דווח חדש על פגמים בתוצאה של SWE-Bench Pro, שמהווה למודל שיטת בדיקה נפוצה. הדווח raises concerns about the reliability and accuracy of evaluating AI models based on this benchmark.
מקור: OpenAI News — לכתבה המלאה
OpenAI פרסם דווח חדש על פגמים בתוצאה של SWE-Bench Pro, שמהווה למודל שיטת בדיקה נפוצה. הדווח raises concerns about the reliability and accuracy of evaluating AI models based on this benchmark.
מקור: OpenAI News — לכתבה המלאה
כתיבת תגובה