Description
Biodiversity monitoring provides the systematic measurements required to assess ecosystem health, predict environmental collapse, and inform global conservation efforts. A key component is the automatic classification of species from image data, where recent deep learning methods have achieved remarkable performance. However, distinguishing visually similar species based on fine-grained characteristics remains challenging when relying solely on visual information. We propose a training-free framework that integrates human knowledge in the form of textual species descriptions using multimodal large language models (MLLMs). We leverage human-curated species descriptions from Wikipedia as an external source of taxonomic knowledge, providing them alongside a query image to an MLLM, which relates visual evidence to textual descriptions to support the classification decision. To accommodate the limited context length of MLLMs, we first narrow the search space by selecting a small set of candidate classes using image features alone. Because the framework requires no additional training and is compatible with different MLLMs, it can be readily applied to new ecological domains while naturally benefiting from ongoing advances in foundation models. We evaluate the proposed knowledge integration approach on three ecological classification tasks: birds, moths, and flowers. In addition to predicting the species label, the MLLM provides a natural-language explanation that relates visual evidence to the retrieved species descriptions, making the classification process inherently interpretable. Our analysis shows that while incorporating textual knowledge often improves classification performance, current limitations of MLLMs, such as hallucinations and self-contradictions, remain evident.
| Topic | Topic 1: : Methods for integrating social and ecological knowledge |
|---|