Preview abstract
Forensic verification often uses a binary “real vs. fake” label that groups fully synthetic, tampered, and AI-retouched images despite their different consequences. We study these modifications through two complementary channels: a camera channel sensitive to capture and processing traces, and a semantic channel capturing scene content. The channels provide continuous evidence rather than deterministic signatures of manipulation history. We instantiate this perspective in 2CAP (2-Channel Authenticity Protocol), pairing a contrastively trained camera encoder with a frozen semantic encoder for (i) reference-free four-class classification through reliability-weighted cross-attention fusion and (ii) reference-based evidence generation. For the latter, query-reference channel similarities and a patch-level saliency map guide a frozen Vision Language Model (VLM) through an Observe–Generate–Refine loop, without forensic instruction tuning of the VLM. On the evaluated benchmark, 2CAP achieves overall AUC .963 and F1 .844; retouching F1 improves from .868 for the strongest compared baseline to .961. Across five VLM configurations, shared-parser paired evaluation shows model-dependent effects: evidence improves change-type accuracy over evidence-free refinement. These results support manipulation-type discrimination while delimiting the benefits of frozen-VLM refinement.View details
Preview abstract
In this paper we present VDTTS, a visual-driven TTS model. Unlike most recent text-to-speech methods which are limited by their lack of ability to generate speech with pauses, emotions, prosody and pitch, is able to do so by taking advantage of an additional silent video as an input.Our method is composed of video and text encoders that are combined via a multi-source attention layer. Speech is generated by a mel-spectrogram decoder followed by a vocoder. We evaluate our method on several challenging benchmarks including VoxCeleb2. To the best of our knowledge this is the first time such a method is trained and evaluated on in-the-wild examples that include unseen speakers.Through a rigorous evaluation we demonstrate the superior performance of our method with respect to other recent work both in terms of objective measures as well as human listening studies.View details
Preview abstract
We present a large-scale dataset for the task of rewriting an ill-formed natural language question to a well-formed one. Our multi-domain question rewriting (MQR) dataset is constructed from human contributed Stack Exchange question edit histories. The dataset contains 427,719 question pairs which come from 303 domains. We provide human annotations for a subset of the dataset as a quality estimate. When moving from ill-formed to well-formed questions, the question quality improves by an average of 45 points across three aspects. We train sequence-to-sequence neural models on the constructed dataset and obtain an improvement of 13.2% in BLEU-4 over baseline methods built from other data resources.View details