You're pledging to donate if the project hits its minimum goal and gets approved. If not, your funds will be returned.
My paper, "Learning Across Layers and Time For Deepfake Detection" has been accepted at ACCV 2026, taking place in Osaka, Japan, this December. We developed a two-stage training framework to use previously underexplored cross-layer patch, register and classification tokens in vision transformers (ViTs) with the proposed Real Video Memory Bank to make video-level reasoning explicitly aware of real-video reference patterns, whose representation space was observed to be more consistent across datasets than deepfake artifacts. We significantly surpassed SOTA baselines on major academic and modern deepfake benchmarks.
I am putting forth this Manifund proposal to accumulate funds to be able to present the paper in-person at the conference in Osaka.
Code: https://github.com/Krishna-Nohwal/Learning-Across-Layers-and-Time-for-Deepfake-Detection
These funds will be used solely towards the presentation of my paper "Learning Across Layers and Time For Deepfake Detection."
I believe that proliferation of AI-generated content is a supersonic tsunami that will leave no field of human thought and endeavour untouched. We must retain our ability to distinguish between what is real and what isn't, but the line is blurring. To counter the wide-ranging negative effects of deepfake media on authenticity, privacy, trust and social discourse, I believe that deepfake detection is an important subfield of AI safety that should not be overlooked.
In this paper, we introduce a two-stage training framework. In Stage 1, we pass [CLS], [REG] and patch tokens from the last four layers of a LoRA-adapted DINOv3-Large into a Layer Token Head, which jointly uses semantic, contextual, and spatial evidence for frame-level classification. Stage 2 uses a temporal transformer with the proposed real-video memory bank to make video-level reasoning explicitly aware of real-video reference patterns, whose representation space is observed to be more consistent across datasets than deepfake videos. Extensive evaluations on academic and modern cross-dataset benchmarks show that my method significantly outperforms existing state-of-the-art deepfake detectors.
Conference Registration : $760, Flights (New Delhi to Osaka) : $680, Hotel for 4 days : $150, Meals for 4 days : $120, Total : $1,710.
I applied for the conference's student travel grant, but it was declined due to the number of applicants and a preference for PhD students. My university is unable to provide financial support at present, and my family cannot cover this without significant financial strain.
I am Krishna Nohwal, a third-year undergraduate in Computer Science & Engineering at Manipal Institute of Technology, India. This paper was the culmination of my work during a summer research
internship at the Indian Institute of Technology Indore, and was co-authored with PhD student Mr. Pawan Soni and Dr. Vivek Kanhangad.
This is the first research paper I've written, and the first one to be accepted. I have continued working in the same domain since the internship, while managing my undergraduate studies as well. I have another paper on deepfake detection under review at WACV (Winter Conference on Applications of Computer Vision, a Core Rank A conference). I'm also working on adapting large pretrained vision transformers for other applications such as medical image segmentation and plant growth modelling.
If this funding fails, I will be unable to present the paper in-person at ACCV. I will also miss out on the oppurtunity to network with fellow researchers and receive feedback on my work.
I have not raised any funds in the last 12 months.