Training-Verifiable and Ownership-Protected Neural Networks

Publication Type:
Thesis
Issue Date:
2026
Full metadata record
As Artificial Intelligence (AI) gains popularity across wide applications, AI models are increasingly shared, outsourced, and commercialized, necessitating robust verification of both provenance (Proof-of-Learning, PoL) and ownership (Proof-of-Ownership, PoO). In existing PoL schemes, verifiers must replay or retrain from the recorded snapshots, causing verification costs to scale with model size. This approaches the full cost of training, making the approach impractical for large or high-capacity models. Fingerprinting and watermarking are two primary methods for protecting the model PoO. However, current watermarking alters model weights and can degrade performance, while fingerprinting often merely verifies uniqueness or requires heavy computation. These limitations become even more prohibitive in collaborative learning such as Split Learning (SL) and Federated Learning (FL), which involve multiple parties and strict privacy requirements. To address the drawback of existing PoL, this thesis proposes PoLO, a unified framework that simultaneously achieves PoL and PoO through chained watermarks. PoLO offers more efficient and privacy-preserving verification than gradient-based PoL methods that risk exposing training data, while providing stronger ownership guarantees than traditional watermark-based PoO methods. To address PoO limitations, a Neural Network Fingerprinting-based Model Authentication Code (NNFMAC) is presented to verify both model uniqueness and ownership without degrading performance. NNFMAC extracts trained model key weights, applies median-based binarization to derive a unique binary fingerprint used as a codebook, and encodes ownership via a newly designed index-based function, producing reliable authentication codes. Furthermore, PoO and label expansion are instantiated for multi-client cooperative SL (CliCooper), where model segments are distributed across training clients to furnish multi-party training verification and copyright attestation while preserving data privacy along the SL pipeline. For FL, a non-interfering fragmented watermarking strategy (FedNIFW) assigns client-specific segments of network layers for embedding, thereby protecting ownership and eliminating cross-client watermark conflicts. The extensive experiments demonstrate that PoLO achieves 99% watermark detection accuracy for ownership verification while preserving data privacy and cutting verification costs to just 1.5–10% of traditional PoL. NNFMAC maintains model accuracy with no additional training overhead and demonstrates strong robustness against a range of attacks. Experiments further show strong privacy protection in SL: input sample reconstruction similarity drops from 0.5 to 0.03, and the PoLO-based instantiation enables secure, efficient multi-party training verification. In FL, client-specific segmented watermarking prevents cross-client conflicts without degrading model accuracy.
Please use this identifier to cite or link to this item: