DeepCodeSeek: Real-Time API Retrieval for Context-Aware Code Generation Paper • 2509.25716 • Published Sep 30, 2025 • 5
Optimizing What Matters: AUC-Driven Learning for Robust Neural Retrieval Paper • 2510.00137 • Published Sep 30, 2025 • 4
Do Enterprise Systems Need Learned World Models? The Importance of Context to Infer Dynamics Paper • 2605.12178 • Published May 12 • 66
EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents Paper • 2605.13841 • Published May 13 • 78
SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding Paper • 2607.10400 • Published Jul 11 • 73
StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments Paper • 2608.24804 • Published 20 days ago • 41
AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling Paper • 2608.26623 • Published 18 days ago • 21
EnterpriseOps-Gym: Environments and Evaluations for Stateful Agentic Planning and Tool Use in Enterprise Settings Paper • 2603.13594 • Published Mar 13 • 150
view article Article AprielGuard: A Guardrail for Safety and Adversarial Robustness in Modern LLM Systems ServiceNow-AI • Dec 23, 2025 • 50
view article Article Apriel-1.6-15b-Thinker: Cost-efficient Frontier Multimodal Performance ServiceNow-AI • Dec 9, 2025 • 84
view article Article SyGra: The One-Stop Framework for Building Data for LLMs and SLMs ServiceNow-AI • Sep 22, 2025 • 14
SyGra: A Unified Graph-Based Framework for Scalable Generation, Quality Tagging, and Management of Synthetic Data Paper • 2508.15432 • Published Aug 21, 2025 • 7