FORTE (Full-cycle Office Real-world Task Evaluation) is a general agent benchmark for evaluating AI agents on daily office productivity across 15 corporate professions.
By chatting or signing in you agree to the Terms and chat-message logging (revocable in History).