- dataset: id: llamaindex/ExtractBench task_id: mean value: 90.29 date: '2026-10-05' source: url: https://huggingface.co/datasets/llamaindex/ExtractBench name: ExtractBench user: SeaWolf-AI notes: 'Pipeline name: darwin_180b_rsi_r3_bf16_vllm_extract_oneshot_structured_output_file_nothink (vllm_extract provider, max_tokens 32768, temperature 0, json_object output, thinking disabled with chat_template_kwargs enable_thinking=false); served checkpoint: FINAL-Bench/Darwin-180B-RSI-R3 on vLLM 0.29.0 with online FP8 quantization, tensor parallel 2. Single run, 370 of 370 documents completed.' - dataset: id: llamaindex/ExtractBench task_id: short value: 95.23 date: '2026-10-05' source: url: https://huggingface.co/datasets/llamaindex/ExtractBench name: ExtractBench user: SeaWolf-AI notes: 'Pipeline name: darwin_180b_rsi_r3_bf16_vllm_extract_oneshot_structured_output_file_nothink (vllm_extract provider, max_tokens 32768, temperature 0, json_object output, thinking disabled with chat_template_kwargs enable_thinking=false); served checkpoint: FINAL-Bench/Darwin-180B-RSI-R3 on vLLM 0.29.0 with online FP8 quantization, tensor parallel 2. Single run, 370 of 370 documents completed.' - dataset: id: llamaindex/ExtractBench task_id: medium value: 88.29 date: '2026-10-05' source: url: https://huggingface.co/datasets/llamaindex/ExtractBench name: ExtractBench user: SeaWolf-AI notes: 'Pipeline name: darwin_180b_rsi_r3_bf16_vllm_extract_oneshot_structured_output_file_nothink (vllm_extract provider, max_tokens 32768, temperature 0, json_object output, thinking disabled with chat_template_kwargs enable_thinking=false); served checkpoint: FINAL-Bench/Darwin-180B-RSI-R3 on vLLM 0.29.0 with online FP8 quantization, tensor parallel 2. Single run, 370 of 370 documents completed.' - dataset: id: llamaindex/ExtractBench task_id: long value: 37.82 date: '2026-10-05' source: url: https://huggingface.co/datasets/llamaindex/ExtractBench name: ExtractBench user: SeaWolf-AI notes: 'Pipeline name: darwin_180b_rsi_r3_bf16_vllm_extract_oneshot_structured_output_file_nothink (vllm_extract provider, max_tokens 32768, temperature 0, json_object output, thinking disabled with chat_template_kwargs enable_thinking=false); served checkpoint: FINAL-Bench/Darwin-180B-RSI-R3 on vLLM 0.29.0 with online FP8 quantization, tensor parallel 2. Single run, 370 of 370 documents completed.'