PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “multimodal artificial intelligence”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Metal-Organic Framework-Based and Metal-Organic Framework-Derived Nanomaterials for Cancer Theranostics and Antibacterial Applications: Advances, Challenges, and Perspectives.

Metal-organic frameworks (MOFs), constructed through coordination-driven self-assembly of metal ions/clusters and organic linkers, have emerged as a uniquely versatile class of porous nanomaterials with broad biomedical potential. Despite substantial clinical progress, both oncological treatment and antimicrobial intervention remain constrained by inadequate tumor-targeting selectivity, multidrug resistance, immunosuppressive tumor microenvironments, and the global proliferation of antibiotic-resistant pathogens, limitations that conventional nanocarrier platforms have addressed only in part. MOF-based and MOF-derived nanomaterials, distinguished by tunable pore architecture, structurally and compositionally adaptable metal nodes, high surface areas, and stimulus-responsive degradability, offer a rational framework for overcoming these barriers. This review systematically examines the synthetic strategies underlying MOF-based and MOF-derived nanomaterials, including pyrolysis, chemical etching, composite modification, and functional group introduction, and their structural determinants of performance. In cancer theranostics, we critically evaluate their roles as multimodal imaging contrast agents, stimulus-responsive drug delivery carriers, and platforms for combination therapies encompassing photodynamic, photothermal, chemodynamic, and immunomodulatory modalities. In antibacterial applications, we analyze the mechanistic basis of MOF-based and MOF-derived activity, including physical membrane disruption, reactive oxygen species-mediated oxidative stress, and sustained metal ion release, alongside strategies targeting biofilm formation and antibiotic resistance. Multifunctional platforms that concurrently integrate cancer theranostic and antibacterial capabilities are further discussed. This review also addresses the principal barriers to clinical translation, encompassing large-scale manufacturing, long-term biosafety, and regulatory approval, and proposes future directions incorporating artificial intelligence-assisted design and materials genomics, underscoring the transformative potential of MOF-based and MOF-derived nanomaterials as next-generation precision nanomedicines. This review establishes a unified mechanistic framework grounded in the intrinsic physicochemical properties of MOF-derived nanomaterials, systematically integrating their applications in cancer theranostics and antibacterial therapy. Critically, it bridges fundamental advances with translational reality by incorporating a rigorous assessment of regulatory pathways, scalable manufacturing constraints, and clinical implementation barriers, and offers a comprehensive, practice-oriented reference for the rational design and responsible translation of MOF-based and MOF-derived nanomaterials.

Theranostic Nanomedicine↗

Feature extraction using recursive cluster-based linear discriminant with application to face recognition.

A novel recursive procedure for extracting discriminant features, termed recursive cluster-based linear discriminant (RCLD), is proposed in this paper. Compared to the traditional Fisher linear discriminant (FLD) and its variations, RCLD has a number of advantages. First of all, it relaxes the constraint on the total number of features that can be extracted. Second, it fully exploits all information available for discrimination. In addition, RCLD is able to cope with multimodal distributions, which overcomes an inherent problem of conventional FLDs, which assumes uni-modal class distributions. Extensive experiments have been carried out on various types of face recognition problems for Yale, Olivetti Research Laboratory, and JAFFE databases to evaluate and compare the performance of the proposed algorithm with other feature extraction methods. The resulting improvement of performances by the new feature extraction scheme is significant.

Algorithms↗

AIM Project A2003: COmputer VIsion in RAdiology (COVIRA).

This paper presents an overview of the COVIRA project, AIM Project No. 2003. The COVIRA consortium is performing research in the area of Multimodality Image Analysis, i.e., Registration and Segmentation. Together with results in the areas of Visualization, User Interface, Digital Anatomy Atlas, Conformal 3D Radiation Therapy Planning, and Cerebral Vessel Tree Reconstruction, clinical validation of initial results is under way at six clinical sites in five European countries. The main objective is to achieve an increase in efficiency and quality of healthcare in Neuro-radiological Diagnosis and Treatment Planning.

Anatomy, Artistic↗

Neuro-fuzzy systems for computer-aided myocardial viability assessment.

This paper describes a multimodality framework for computer-aided myocardial viability assessment based on neuro-fuzzy techniques. The proposed approach distinguishes two main levels: the modality-independent inference level and the modality-dependent application level. This two-level distinction releases the hard constraint of multimodality image registration. An abstract description template is used to describe the different myocardial functions (contractile function, perfusion, metabolism). Parameters extracted from different image modalities are combined to derive a diagnostic image. The neuro-fuzzy techniques make our system transparent, adaptive and easily extendable. Its effectiveness and robustness are demonstrated in a positron emission tomography/magnetic resonance imaging data fusion application.

Artificial Intelligence↗

Ensembling local learners through multimodal perturbation.

Ensemble learning algorithms train multiple component learners and then combine their predictions. In order to generate a strong ensemble, the component learners should be with high accuracy as well as high diversity. A popularly used scheme in generating accurate but diverse component learners is to perturb the training data with resampling methods, such as the bootstrap sampling used in bagging. However, such a scheme is not very effective on local learners such as nearest-neighbor classifiers because a slight change in training data can hardly result in local learners with big differences. In this paper, a new ensemble algorithm named Filtered Attribute Subspace based Bagging with Injected Randomness (FASBIR) is proposed for building ensembles of local learners, which utilizes multimodal perturbation to help generate accurate but diverse component learners. In detail, FASBIR employs the perturbation on the training data with bootstrap sampling, the perturbation on the input attributes with attribute filtering and attribute subspace selection, and the perturbation on the learning parameters with randomly configured distance metrics. A large empirical study shows that FASBIR is effective in building ensembles of nearest-neighbor classifiers, whose performance is better than that of many other ensemble algorithms.

Algorithms↗

AI-Supported, Integrative Prediction of Postoperative Delirium: Protocol for the CONFUSED Study.

BACKGROUND: Postoperative delirium (POD) is a frequent and serious complication in older surgical patients, characterized by acute cognitive dysfunction and fluctuating levels of consciousness. POD is associated with prolonged hospitalization, long-term cognitive decline, reduced quality of life, and increased mortality. Despite its clinical relevance, the underlying pathophysiological mechanisms remain poorly understood, and reliable biomarkers for early prediction and prevention are lacking. OBJECTIVE: The CONFUSED study aims to identify molecular and clinical predictors of POD by integrating clinical data with proteomic, transcriptomic, and epigenetic analyses. The primary objective is to develop predictive models for POD using multimodal data. Secondary objectives include the identification of delirium-associated genes, proteins, and epigenetic signatures, as well as the exploration of patient subgroups at increased risk for POD. METHODS: CONFUSED is a prospective observational cohort study conducted at a German university hospital. Adult patients undergoing major surgery under general anesthesia will be enrolled until 100 cases of POD have been observed, which is expected to require a total sample size of approximately 200 to 300 patients. Blood samples are collected at 4 predefined time points: before premedication, immediately after surgery, and on postoperative days 2 and 5. Samples undergo comprehensive proteomic profiling, transcriptomic analysis using RNA microarrays, DNA methylation analysis, and genotyping of selected polymorphisms. Clinical data, including demographics, comorbidities, perioperative variables, medications, and delirium assessments using the Confusion Assessment Method (CAM) and CAM for the intensive care unit, are systematically recorded. Statistical analyses include univariate and multivariate methods, as well as machine learning approaches such as random forests and support vector machines, to identify relevant biomarkers and develop predictive models. The study protocol follows STROBE (Strengthening the Reporting of Observational Studies in Epidemiology) and TRIPOD (Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis) guidelines and was approved by the responsible ethics committees. RESULTS: The study was registered in the German Clinical Trials Register (DRKS00033854) on March 18, 2024. Recruitment started in January 2024 and is ongoing at the time of manuscript submission. As of now, 135 patients have been enrolled. Sample collection and laboratory analyses are ongoing. Data analysis began in January 2026, with first results anticipated in July 2026. Final data lock is anticipated after the completion of recruitment. CONCLUSIONS: By integrating multimodal molecular data with clinical parameters and applying advanced machine learning techniques, the CONFUSED study aims to improve the prediction and understanding of POD. The results are expected to support the development of personalized preventive strategies and contribute to improved perioperative care for patients at risk of POD.

Humans↗

Large language models in bioinformatics: a comprehensive survey.

The emergence of foundation models with trillion-level parameters has redefined the landscape of artificial intelligence. Various fields are developing their own large-scale models, which can solve many problems within the field and improve work efficiency. Biological large-scale models are a cross-disciplinary research field that combines mathematics, computer science, and biology, aiming to simulate and understand the structure, function, and dynamic changes of biological systems through the establishment of complex computational models. This field covers multiple levels such as biological pathways, population dynamics, protein folding, etc., providing us with tools for deep exploration of the mysteries of life and applications in medicine, ecology, and other fields. This article reviews the background and research status of biological large-scale models, and discusses future directions. Large language models (LLMs) and other large-scale foundation models have rapidly advanced in recent years, enabling powerful representation learning and generation across text, sequences, and multimodal data. In bioinformatics and biomedicine, these models are increasingly used to analyze genomic sequences, infer protein properties and structures, support drug discovery, and integrate heterogeneous biomedical evidence. This survey reviews the basic principles of LLMs and summarizes representative applications in (i) gene and genome sequence analysis, (ii) protein structure and function prediction, and (iii) drug design, including virtual screening and personalized medicine. We also discuss emerging multi-model modeling approaches, as well as key challenges such as data quality and privacy, interpretability, generalization to new organisms and tasks, and responsible deployment in health-related settings. Finally, we outline future directions for developing reliable, scalable, and explainable bioinformatics foundation models.

bioinformatics↗

ICOHR: intelligent computer based oral health record.

The majority of work on computer use in the dental field has focused on non-clinical practice management information needs. Very few computer-based dental information systems provide management support of the clinical care process, particularly with respect to quality management. Traditional quality assurance methods rely on the paper record and provide only retrospective analysis. Today, proactive quality management initiatives are on the rise. Computer-based dental information systems are being integrated into the care environment, actively providing decision support as patient care is being delivered. These new systems emphasize assessment and improvement of patient care at the time of treatment, thus building internal quality management into the caregiving process. The integration of real time quality management and patient care will be expedited by the introduction of an information system architecture that emulates the gathering and storage of clinical care data currently provided by the paper record. As a proposed solution to the problems associated with existing dental record systems, the computer-based patient record has emerged as a possible alternative to the paper dental record. The Institute of Medicine (IOM) recently conducted a study on improving the efficiency and accuracy of patient record keeping. As a result of this study, the IOM advocates the development and implementation of computer-based patient records as the standard for all patient care records. This project represents the ongoing efforts of The University of Iowa College of Dentistry's collaboration with the University of Uppsala Data Center, Uppsala, Sweden, on a computer-based patient dental record model. ICOHR (Intelligent Computer Based Oral Health Record) is an information system which brings together five important parts of the patient's dental record: medical and dental history; oral status; treatment planning; progress notes; and a Patient Care Database, generated from their clinical care information (the database is also stored in the ICOHR). ICOHR is designed to be integrated into a traditional practice management system. The components of the ICOHR system support the use of various types of clinical care quality management tools, including medical alerts, clinical care guidelines, care modifiers, and diagnostic decision support. Data input is multimodal, so the user may use both voice recognition and direct input with a digitizer board to enter information into the database. ICOHR is designed to be integrated into the clinical environment in an ergonomic fashion in order to facilitate the unobtrusive and efficient acquisition of patient information. ICOHR is currently under clinical evaluation in both private practice and institutional environments. The private practice is a large general dentistry practice with over twenty sites scattered throughout a large metropolitan area. The institutional settings are a College of Dentistry and a Hospital Dentistry program. The evaluations have started in two of the sites and the other site will be phased in during the next six months. Our demonstration of the system will include both prepared presentations of the system's various functions and provide an opportunity for hands-on use of the system for interested attendees.

Artificial Intelligence↗

Generalized RLS approach to the training of neural networks.

Recursive least square (RLS) is an efficient approach to neural network training. However, in the classical RLS algorithm, there is no explicit decay in the energy function. This will lead to an unsatisfactory generalization ability for the trained networks. In this paper, we propose a generalized RLS (GRLS) model which includes a general decay term in the energy function for the training of feedforward neural networks. In particular, four different weight decay functions, namely, the quadratic weight decay, the constant weight decay and the newly proposed multimodal and quartic weight decay are discussed. By using the GRLS approach, not only the generalization ability of the trained networks is significantly improved but more unnecessary weights are pruned to obtain a compact network. Furthermore, the computational complexity of the GRLS remains the same as that of the standard RLS algorithm. The advantages and tradeoffs of using different decay functions are analyzed and then demonstrated with examples. Simulation results show that our approach is able to meet the design goals: improving the generalization ability of the trained network while getting a compact network.

Algorithms↗

A refined algorithm for multisensor image registration based on pixel migration.

Multimodality image registration via pixel migration is a powerful approach. However, it suffers from a serious problem--the global maximum on the sum of squared gradient magnitude (SSG) surface does not correspond to the correct solution of registration. To solve the problem, we partition the search space into feasible and infeasible regions. The genetic algorithm (global optimizer) is used to obtain a good initial estimate of registration parameters and followed by a fast refining with Powell's approach (local optimizer). The experimental results demonstrate that the use of this modified pixel migration algorithm on multisensor image registration is very effective.

Algorithms↗

Machine learning for detection and diagnosis of disease.

Machine learning offers a principled approach for developing sophisticated, automatic, and objective algorithms for analysis of high-dimensional and multimodal biomedical data. This review focuses on several advances in the state of the art that have shown promise in improving detection, diagnosis, and therapeutic monitoring of disease. Key in the advancement has been the development of a more in-depth understanding and theoretical analysis of critical issues related to algorithmic construction and learning theory. These include trade-offs for maximizing generalization performance, use of physically realistic constraints, and incorporation of prior knowledge and uncertainty. The review describes recent developments in machine learning, focusing on supervised and unsupervised linear methods and Bayesian inference, which have made significant impacts in the detection and diagnosis of disease in biomedicine. We describe the different methodologies and, for each, provide examples of their application to specific domains in biomedical diagnostics.

Algorithms↗

Performance enhancement for audio-visual speaker identification using dynamic facial muscle model.

Science of human identification using physiological characteristics or biometry has been of great concern in security systems. However, robust multimodal identification systems based on audio-visual information has not been thoroughly investigated yet. Therefore, the aim of this work to propose a model-based feature extraction method which employs physiological characteristics of facial muscles producing lip movements. This approach adopts the intrinsic properties of muscles such as viscosity, elasticity, and mass which are extracted from the dynamic lip model. These parameters are exclusively dependent on the neuro-muscular properties of speaker; consequently, imitation of valid speakers could be reduced to a large extent. These parameters are applied to a hidden Markov model (HMM) audio-visual identification system. In this work, a combination of audio and video features has been employed by adopting a multistream pseudo-synchronized HMM training method. Noise robust audio features such as Mel-frequency cepstral coefficients (MFCC), spectral subtraction (SS), and relative spectra perceptual linear prediction (J-RASTA-PLP) have been used to evaluate the performance of the multimodal system once efficient audio feature extraction methods have been utilized. The superior performance of the proposed system is demonstrated on a large multispeaker database of continuously spoken digits, along with a sentence that is phonetically rich. To evaluate the robustness of algorithms, some experiments were performed on genetically identical twins. Furthermore, changes in speaker voice were simulated with drug inhalation tests. In 3 dB signal to noise ratio (SNR), the dynamic muscle model improved the identification rate of the audio-visual system from 91 to 98%. Results on identical twins revealed that there was an apparent improvement on the performance for the dynamic muscle model-based system, in which the identification rate of the audio-visual system was enhanced from 87 to 96%.

Adult↗

Experiments with repeating weighted boosting search for optimization in signal processing applications.

Many signal processing applications pose optimization problems with multimodal and nonsmooth cost functions. Gradient methods are ineffective in these situations, and optimization methods that require no gradient and can achieve a global optimal solution are highly desired to tackle these difficult problems. The paper proposes a guided global search optimization technique, referred to as the repeated weighted boosting search. The proposed optimization algorithm is extremely simple and easy to implement, involving a minimum programming effort. Heuristic explanation is given for the global search capability of this technique. Comparison is made with the two better known and widely used guided global search techniques, known as the genetic algorithm and adaptive simulated annealing, in terms of the requirements for algorithmic parameter tuning. The effectiveness of the proposed algorithm as a global optimizer are investigated through several application examples.

Algorithms↗

Category learning through multimodality sensing.

Humans and other animals learn to form complex categories without receiving a target output, or teaching signal, with each input pattern. In contrast, most computer algorithms that emulate such performance assume the brain is provided with the correct output at the neuronal level or require grossly unphysiological methods of information propagation. Natural environments do not contain explicit labeling signals, but they do contain important information in the form of temporal correlations between sensations to different sensory modalities, and humans are affected by this correlational structure (Howells, 1944; McGurk & MacDonald, 1976; MacDonald & McGurk, 1978; Zellner & Kautz, 1990; Durgin & Proffitt, 1996). In this article we describe a simple, unsupervised neural network algorithm that also uses this natural structure. Using only the co-occurring patterns of lip motion and sound signals from a human speaker, the network learns separate visual and auditory speech classifiers that perform comparably to supervised networks.

Algorithms↗

A biometric identification system based on eigenpalm and eigenfinger features.

This paper presents a multimodal biometric identification system based on the features of the human hand. We describe a new biometric approach to personal identification using eigenfinger and eigenpalm features, with fusion applied at the matching-score level. The identification process can be divided into the following phases: capturing the image; preprocessing; extracting and normalizing the palm and strip-like finger subimages; extracting the eigenpalm and eigenfinger features based on the K-L transform; matching and fusion; and, finally, a decision based on the (k, l)-NN classifier and thresholding. The system was tested on a database of 237 people (1,820 hand images). The experimental results showed the effectiveness of the system in terms of the recognition rate (100 percent), the equal error rate (EER = 0.58 percent), and the total error rate (TER = 0.72 percent).

Algorithms↗

Brain-computer interaction research at the Computer Vision and Multimedia Laboratory, University of Geneva.

This paper describes the work being conducted in the domain of brain-computer interaction (BCI) at the Multimodal Interaction Group, Computer Vision and Multimedia Laboratory, University of Geneva, Geneva, Switzerland. The application focus of this work is on multimodal interaction rather than on rehabilitation, that is how to augment classical interaction by means of physiological measurements. Three main research topics are addressed. The first one concerns the more general problem of brain source activity recognition from EEGs. In contrast with classical deterministic approaches, we studied iterative robust stochastic based reconstruction procedures modeling source and noise statistics, to overcome known limitations of current techniques. We also developed procedures for optimal electroencephalogram (EEG) sensor system design in terms of placement and number of electrodes. The second topic is the study of BCI protocols and performance from an information-theoretic point of view. Various information rate measurements have been compared for assessing BCI abilities. The third research topic concerns the use of EEG and other physiological signals for assessing a user's emotional status.

Animals↗

Bayesian modeling of dynamic scenes for object detection.

Accurate detection of moving objects is an important precursor to stable tracking or recognition. In this paper, we present an object detection scheme that has three innovations over existing approaches. First, the model of the intensities of image pixels as independent random variables is challenged and it is asserted that useful correlation exists in intensities of spatially proximal pixels. This correlation is exploited to sustain high levels of detection accuracy in the presence of dynamic backgrounds. By using a nonparametric density estimation method over a joint domain-range representation of image pixels, multimodal spatial uncertainties and complex dependencies between the domain (location) and range (color) are directly modeled. We propose a model of the background as a single probability density. Second, temporal persistence is proposed as a detection criterion. Unlike previous approaches to object detection which detect objects by building adaptive models of the background, the foreground is modeled to augment the detection of objects (without explicit tracking) since objects detected in the preceding frame contain substantial evidence for detection in the current frame. Finally, the background and foreground models are used competitively in a MAP-MRF decision framework, stressing spatial context as a condition of detecting interesting objects and the posterior function is maximized efficiently by finding the minimum cut of a capacitated graph. Experimental validation of the proposed method is performed and presented on a diverse set of dynamic scenes.

Algorithms↗

Design methods and architectural issues of integrated medical image data base systems.

The past 20 years have seen tremendous changes in medical imaging techniques. New modalities and protocols are expanding the available digital image data at a rapid rate. Yet a framework for gathering, managing, and using multimodal image information is an integrated database environment is missing. The purpose of this paper is to present the experience of implementing an integrated medical image database system at UCSF. We discuss the general system architecture, software design methods, and specific database tools and illustrate them with application examples. Two immediate issues conforming the building of medical image database systems are: lack of supporting infrastructure and inability to index images by contest. To circumvent these problems, the evolutionary medical image database system being implemented at UCSF is based on a three-tiered client-server architecture: client medical workstations, database application servers, and a hospital-integrated picture archiving and communication system (HIP-PACS). The approach used to integrate content-based retrieval and knowledge base techniques within the existing HI-PACS to make the whole database system useful in medicine.

Adult↗