A typical example includes various forms of 'Margot Robbie', and fine-tuning models/LoRas created on her pictures. The goal is straightforward but daunting: clean, standardize, extract meaningful keywords, and cluster these terms based on their relational context to each other.
Given the sheer volume of data, manual sorting is impractical. Hence, I'm looking for scalable, automated solutions.
I'm here to gather insights on tools, libraries, and algorithms that might be particularly effective for this kind of task. If anyone has tackled similar challenges or has relevant experience, your advice would be invaluable.
Appreciate any pointers!