Granite Vision 4.1 4B read the scans, then failed the job

Granite Vision 4.1 4B handled individual synthetic invoice scans better than expected, but did not complete the paperwork audit. The available local multi-image path failed, and the pipeline workaround still broke at final JSON and proof-code closure.

Granite Vision 4.1 4B read the scans, then failed the job

Granite Vision 4.1 4B is positioned as a small open vision-language model for enterprise document extraction. That is close enough to Local Model Bench territory to be interesting: invoice scans, structured fields, source evidence, and final audit artifacts.

The result was split. The model extracted visible facts from individual scans well. But the benchmark is not an OCR demo. The task ends only when the model produces a valid final audit JSON with correct evidence and a computed proof code. Granite did not get there.

Where it worked

  • Extracted visible fields from individual invoice scans surprisingly well.
  • Correctly recognized the quote as not an invoice.
  • Found important visual annotations including short payment and under-review status.

Where it failed

  • The direct multi-image local run did not complete in the available runtime.
  • The completed pipeline workaround produced invalid JSON.
  • The proof code was left as a formula and calculated from the wrong values.
  • That makes it unsuitable as an end-to-end paperwork agent in this setup.