What ended up working - a combo of four layers. OCR pulls button labels, headings, body copy from every image. On top of that, a vision model with my own prompt dictionary for UI patterns: modal, toast, tabs, empty state, chart, dashboard. Wrote the dictionary myself, from what I actually look for. Plus manual tags users drop in two clicks. And fuzzy search over filenames as a fallback.