Multimodality has the potential to facilitate richer interaction styles in both information retrieval and learning environments. However, its true potential will not be realised unless consideration is given to the application of combined modalities. This paper asserts that multimedia output from a system actually requires multimodality on the part of the user in order to ensure that the effectiveness of the communication or information is not lost. The notion of a “multi-modelling” approach to interaction along with the use of gestures and metaphors have been examined and two systems are described which attempt to implement these approaches.