Three of twelve models named Stampli first on the direct prompt; two named BILL. Stampli was named by eleven of the twelve models and BILL by twelve and Stampli carries 37 labels and BILL 52, so the shares are not directly comparable.
Named in eight categories this edition.
Named in twelve categories this edition.
Share is the count of first choices across the direct, paraphrase, budget and scale prompts over all twelve models, for a mid-market B2B company; rank is within the category; every quote names the model and the prompt it came from. Both figures come from the accounts payable automation page.
Bold names in an answer are the products the judge labeled a first choice; a model naming several gives each of them that label. The full answer text for every row is in the record.
Every negative label with a quote, up to three, then the highest-weighted positives, up to three. Two of two in this category shown.
“I'd start with Stampli. It's the strongest "classic AP automation" choice for mid-market companies” GPT-5.4 mini · direct prompt · first choice
“Stampli (best for ERP integration and conversational AI that actually works on messy invoices)” DeepSeek V4 Flash · negative prompt · first choice
Every negative label with a quote, up to three, then the highest-weighted positives, up to three. Two of two in this category shown.
“Verdict: Avoid if you have complex approval workflows, rely on tight cash flow, or have a low tolerance for support friction.” DeepSeek V4 Flash · negative prompt · hard negative
“BILL AP/AR and AvidXchange are the top recommendations due to their strong user ratings, scalability, and comprehensive feature sets.” Mistral Small · paraphrase prompt · first choice
Comparisons are drawn for the top three products in each category. The output is the models' output; nothing here is a recommendation by the index.