Skylume audience clustering experiment by Allan BenameurSkylume audience clustering experiment by Allan Benameur

Skylume audience clustering experiment

Allan Benameur

Allan Benameur

Goal
I built a Python pipeline to group a Bluesky audience from profile biographies. The aim was to find useful audience segments without relying on manually maintained keyword lists.
Approach
The pipeline generated biography embeddings with Sentence Transformers, reduced dimensionality with UMAP and clustered profiles with HDBSCAN. LDA and KeyBERT were used to inspect recurring themes, with spaCy supporting text processing.
Result
The pipeline ran successfully and some groups were visible, but most profiles remained weakly differentiated. It was also too slow for a real-time product path. The chart shows 439 profiles, including three detected groups and 334 profiles that did not belong to a useful cluster.
Decision
The result was useful as an experiment, not as a product feature. I replaced it with simpler TypeScript keyword heuristics that were faster and easier to explain. The Python pipeline remains archived in the repository and was never deployed to production.
Like this project

Posted Sep 14, 2026

Python pipeline for semantic audience clustering with embeddings, UMAP and HDBSCAN. Archived after latency and cluster quality failed product needs.