Problem Statement

The OvaHimba in northern Namibia and the San in Southern Africa are indigenous peoples, with the San – as ‘hunters and gatherers’ – belonging to the Khoisan group alongside the Khoikhoi.

  • What problems arise from a lack of representation in AI image generation?
  • What ethical and technical aspects are important?
  • Why does AI fail to accurately depict the Himba and the San?
  • How efficiently can diffusion models be trained using one’s own resources

Idea and Concept

Basic idea: Closing representation gaps through a specialised image-generation model for under-represented groups (OvaHimba and San).

Approach: The creation of a bespoke image dataset and the fine-tuning of an AI model using this dataset enable the generation of authentic representations through targeted prompts.

Implementation

The project employs a structured approach: initially, the Stable Diffusion model was used as a basis. Training datasets were created using Kohya, utilising the captioning function (WD14) and the TrainLoRA module. These data were used to create specialised models for the Himba and San peoples, taking gender into account.

Flux was introduced into the project: a model that delivers better results through flow matching. Here, too, text-image pairs were created to generate high-quality visual content. Finally, the prompts were optimised and the generated images were post-processed. Technologies used:

  • Image generation: Stable Diffusion (v1.5, 3XL), Flux, Kohya
  • Operating system and environment: Linux, Python 3.12.7
  • Image editing: Adobe Photoshop, Lightroom
  • Other tools: Booru Tag Manager, CUDA, JoyCaption, TagGui

🛈 Photo 1-3: Original photos of the OvaHimba people.
🛈 Photo 4: AI-generated image before training the model.
🛈 Photo 5: AI-generated image after training the model.

Module

Bachelor Software- and Multimediaproject

Duration

06/2024 – 03/2025

Team Member/s

Abdelrahman Abdelazim
Dominik Lubos
Mehmet Ubeyd Yildiz
Tobias Rogowski
Nicole Vögele