AI & ML interests
None defined yet.
Recent Activity
Papers
Beyond Next-Token Prediction: An RLVR Proof of Concept for Tool-Use Agents on Atlassian Workflows
World Feedback for Clinical Agents: Diagnosing RL in FHIR Environments
Personal Assistant Benchmark
Scores a personal assistant by what it did on the device
Healthcare Document-Grounded QA Benchmark (Extending GDP.pdf)
Document-work benchmark for healthcare persona
JE Validation Console
Evaluate AI models on journal entry audit tasks
RLEaaS RL-ADA Arena
Co-evolutionary adversarial training demo (DA vs CA)
Atlassian Cloud RL Β· Trajectory Playground
Explore and compare RL task trajectories
BA Agent RL Environment and Benchmark
RL env & benchmark for enterprise BA agents
A2A Marketplace
Cached replays of 140 agent-to-agent negotiation rollouts
Sales & RevenueOps Agent RL Environment
RL environment for sales & revenue-ops agents
EHR Clinical-Task Agent RL Environment and Benchmark
RL environment & benchmark for clinical EHR agents
PersonalAssistantBench β iOS Assistant RL
Run and evaluate simulated iPhone assistant tasks
Ad Studio Agent
Generate a personalized ad and receive a quality score
ALE - Trust Layer evaluation
Assess AI benchmark scores with a multiβdimensional trust report
MedMosaic: A Challenging Large Scale Benchmark of Diverse Medical Audio
Interactive demo for the MedMosaic medical-audio benchmark