Two of twelve models named Upflow first on the direct prompt; zero named Gaviti. Upflow was named by seven of the twelve models and Gaviti by eight and both carry 13 labels, so the shares below are directly comparable.
Named in five categories this edition.
Named in four categories this edition.
Share is the count of first choices across the direct, paraphrase, budget and scale prompts over all twelve models, for a mid-market B2B company; rank is within the category; every quote names the model and the prompt it came from. Both figures come from the cash application page.
Across every category in the September 2026 Edition, Upflow and Gaviti were named in the same answer twenty-five times, of the 121 answers naming Upflow and the 93 naming Gaviti. In those answers Gaviti took the first choice four times and Upflow five.
| Model | Direct | Paraphrase | Comparative | Budget-constrained | Scale-constrained | Negative |
|---|---|---|---|---|---|---|
| Claude Haiku 4.5 | ||||||
| GPT-5.4 mini | ||||||
| Gemini 3.5 Flash | ||||||
| Perplexity Sonar | ||||||
| Grok 4.1 Fast | ||||||
| Mistral Small | ||||||
| DeepSeek V4 Flash | ||||||
| Llama 4 Maverick | ||||||
| Qwen 3.7 Flash | ||||||
| Kimi K2 | ||||||
| GLM 4.7 FlashX | ||||||
| MiniMax M2.5 |
Bold names in an answer are the products the judge labeled a first choice; a model naming several gives each of them that label. The full answer text for every row is in the record.
Every negative label with a quote, up to three, then the highest-weighted positives, up to three. Two of two in this category shown.
“Upflow - Good fit for mid-sized and scaling companies (typically $10M-$500M revenue) that need cash application to feed directly into AR and collections” Llama 4 Maverick · direct prompt · first choice
“Upflow is often recommended as the best overall solution due to its seamless integration with collections and AR workflows” Mistral Small · direct prompt · first choice
Every negative label with a quote, up to three, then the highest-weighted positives, up to three. Two of two in this category shown.
“lowest invoice-matching accuracy score (6.9/10) among major peers. If matching is your core problem, look elsewhere.” DeepSeek V4 Flash · paraphrase prompt · soft negative
“AI-powered cash application, but may be overkill for very small budgets” MiniMax M2.5 · budget prompt · soft negative
Comparisons are drawn for the top eight products in each category, each against each. The output is the models' output; nothing here is a recommendation by the index.