Train more specialised models

#180 · open · 0 comments

View on GitHub ↗

LightArrowsEXE

The current models have seen lots of use, and seem to be generally good for their intended jobs (~~so people really need to stop using it for unintended stuff~~). But the current models are limited to specific problems, and specific values. If anyone has any other neat ideas, feel free to comment them here too. ## Current models - [ ] **Extend the list of delowpassing models**, including both a "regular" version, and a "dempeg2" version (gone through an mpeg2 encoder) for use with DVDs. The dempeg2 versions could be limited to horizontal delowpassing models only (as they're intended for DVDs). - [ ] **Extend the dempeg2 models**. Ideally, I'd have "levels" of compression to help users control the strength a bit better. This could also be improved via some kind of masking logic a la DPIR. ## New models I want to also introduce new models to deal with specific kinds of artefacting that regular filtering has trouble with. - [ ] **De-HDCAM**: not too dissimilar from the previous delowpassing models, but way more targeted. This depends on us figuring out the exact resampling done during HDCAM, as well as see how it may differ between authoring companies. I know for example most of J.C. Staff's authoring company's HDCAM appears to differ from others, and this might extend to other authoring houses too. The ideal scenario would involve somehow acquiring one of those HDCAM machines so we can reverse engineer them, or if all else fails, literally run clean video through that and use that as the LQ sample. - [ ] **HD-downscale dehalo resampling**: A lot of shows were HD productions but only have DVDs available (or very bad WEB streams). A while ago we uncovered the exact resample kernel and values used by a lot of authoring houses for this process, but by virtue of it being a downscale, we can't really "fix" it through regular means. This model should try to help with that, and try to make it closely match more reasonable downscaling methods. This can be split into multiple models using different downscalers as GT. - [ ] **De-unsharpening**: A lot of shows are getting post-sharpened again as of late, so descaling them is impossible, and dehaloing only fixes up some of the damage while also being very destructive if you have extreme sharpening cases. I plan to have at least two models for this, one that *only* fixes the unsharpening, and one that includes a rescaled GT to help further improve lineart, intended for shows that are clearly *not* native FHD, but can't be rescaled. This relies on us figuring out what kind of unsharpening settings are common among studios (since I assume this is done at the studio level). - [ ] **Per-streaming-service de-AVC**: Similar to dempeg2, but more targeted at specific settings based on streaming services. Would be a lot more conservative ideally, and try to lightly denoise at most. - [ ] **De-point resized chroma**: We have some current methods that work (ChromaReconstruct and ArtCNN DN (dev?)), but it's probably not a bad idea to train a model specifically on 422 -> 420 point-downscaled chroma. This can become a bit of a problem however, in that I need to actually, like, get 422 material to train on somehow. - [ ] **De-"chroma shift"**: For some reason, a number of BDs randomly have a chroma shift to the left. This blurs the chroma, and thus requires some abusing of the `vsdenoise.frequency_merge` function to "undo" (using un-shifted chroma as a reference, like from a WEB source). If this were simply mistagged chroma, that'd be easy to deal with (just set the correct chroma location), but it being shifted means you can't fix it as easily. Ideally we should shift chroma ourselves on a clean source with multiple known-shifting kernels and distance in px, and focus heavily on not overshooting. ## POC - [ ] **Field-inpainting/deinterlacing**: What it says on the tin. Might be better left to someone like Artorius, but in the absence of that, I can try to experiment somewhat. - [ ] **Linear-light rescale**: Since you can't really descale linear light sources, I want to see if I can train some models to help with that, intended entirely as a POC. Test case would be New Game!, and if that proves to be useful, maybe I'll extend it and make it a more general models. - [ ] **"Fake rescaling"**: Train models on properly rescaled anime, and see if that helps with dealing with anime that can't be descaled (reliably). I can't imagine this going well, but hey, maybe it'll work... ## ML-approaches to current manual filtering Currently, a lot of time is spent on certain manual operations that could _theoretically_ maybe be shortened/semi-automated using ML-based approaches, with the caveat that this requires highly specialised datasets, APIs, and workflow adjustments (which may include tooling changing to accommodate them) to be introduced to the ecosystem. Below are a handful of ideas I had when yapping in a certain [Discord channel](https://discord.com/channels/1132160355829821561/1132160361311764523/1528756873707262112) of what some of these could do: - [ ] Determining the most likely upscale kernel used over the course of multiple episodes (including checking for likely switches and different combinations of resampling params, and possibly also finding previously rare/previously unencountered kernels (such as Kaiser for Soul Eater). - [ ] Improved automatic IVTC metrics and field matching, including when to determine something must be deinterlaced (and letting the user pass a callable to handle that such as QTGMC). This would mostly be used as a helper during regular manual IVTC such as the [Wobbly workflow](https://wobbly.encode.moe/) and necessitate changes to Wobbly to allow for this, or for it to be taken into consideration when designing a replacement (such as the proposed [vs-view Wobbly plugin](https://github.com/Jaded-Encoding-Thaumaturgy/vs-view/issues/41) should that ever take off). - [ ] Nice to have: automatic crossfade detection and handling - [ ] Nice to have: automatic (cross-field-)blending detection and handling - [ ] Nice to have: per-plane handling using the (IVTC'd) luma as a ground truth reference (mainly useful for sources such as the [Persona 3 OP](https://www.youtube.com/watch?v=XxNAwZ-A88w), [CANAAN](https://anidb.net/anime/6275), and a handful of other shows with desynced chroma). - [ ] Determining whether lowpassing was applied to a video during authoring using a presumed "clean" reference source, and determining which params were likely used while trying to account for compression. Bonus points if this also could handle sources with multiple distinct lowpass parameters applied (such as the [Fate/strange Fake](https://anidb.net/anime/18093) BDs) - [ ] Determining which sharpening filters were run over a video and with which params, given a reference table of common sharpeners used in NLEs and other known production chains Important to note is that these would be _companion pieces_ to help determine whether a source requires specialised handling or whether certain filtering approaches (including the other proposed models) should be used in the first place, as well as what kind of models may need to be trained (such as for example in the delowpass case). --- --- Besides this, I also still want to figure out the best way to avoid resampling artefacting. Most ideal would be passing numpy arrays and reconstructing YUV frames through some kind of magic system, but in the absence of that, I'll have to figure out another way. Plane stacking is one such idea, but I'm not sure how reliable that truly is. Should probably discuss this with Art once I get here. I should probably also use chromaloc-centre from now on, oopsie. I also want to, if possible, extend that into a VS library to make it easier for myself and others to train material via VS in the future, but that's in the far-off future. ---

Comments