@rygorousi
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- Web
Account-level information from X, not a live location or the device used for a specific post.
Abstraction maker, abstraction breaker. @rygorous@mastodon.gamedev.place he/him
Joined December 2009
- Tweets83.6K
- Following91
- Followers14.8K
- Likes4K
New blog post: "UNORM and SNORM to float, hardware edition" fgiesen.wordpress.com/2024/1…
We just released Oodle 2.9.13.
Significantly increased BC7 encoding speed (about 20-25% encode time reduction for non-RDO on typical content, 25-30% encode time reduction for RDO) at slightly increased quality.
Also several bug fixes and experimental WASM 64-bit support.
New blog post: "Exact UNORM8 to float" fgiesen.wordpress.com/2024/1… a satisfying solution to a problem that, quite possibly, nobody has
New blog post: "BC7 optimal solid-color blocks" fgiesen.wordpress.com/2024/1… clearing out my "I should write this up" queue, this technique is from... *checks git logs* May 2017. Oh my. (I have quite the backlog.)
New blog post: "Why those particular integer multiplies?" fgiesen.wordpress.com/2024/1… some explanation and some speculation on the integer SIMD multiplies offered in x86, along with some history
New blog post: "Inserting a 0 bit in the middle of a value" fgiesen.wordpress.com/2024/1… I guess it's 2-for-1 bit hacks week.
New blog post: "Zero or sign extend" fgiesen.wordpress.com/2024/1…
We released Oodle 2.9.12 last week:
radgametools.com/oodlehist.h…
Some SDK/compiler updates and bug fixes. Also, max texture size limit bumped from 16384x16384 to 2097152x2097152, which should be good for at least the next 4 months or so.
New blog post: "Entropy decoding in Oodle Data: x86-64 6-stream Huffman decoders" fgiesen.wordpress.com/2023/1…
We just released Oodle 2.9.11 (website isn't updated yet, soon!)
This one is focused on Oodle Texture improvements.
* Faster end-to-end latency in multi-threaded encoding especially on 24+ core machines, most noticeable for BC[145].
* Re-designed mode/partition selection logic in baseline BC7 encoder (no RDO). Roughly 2x faster at all encoding effort levels at typically same or better result quality. For BC7 RDO, works out to around a 1.2x speed-up typically.
Oodle Texture PSA: if you're on a Zen 4 machine, in current releases, encode textures with OodleTex_BCNFlag_AvoidWideVectors (disables usage of AVX-512 instructions).
Some of the hot AVX-512 loops heavily use the store forms of VPCOMPRESSD which are quite slow on Zen 4.
Fabian Giesen retweeted
The world has dimmed for us.
With sadness and his loved ones in our hearts, we say farewell to our friend and fellow main orga - @acrydfr. More than 25 years of demoparties wouldn't have been the same without him. And will never be again. We miss you, Ben.
pouet.net/topic.php?which=12…
New blog post: "A very brief BitKnit retrospective" fgiesen.wordpress.com/2023/0… Small codec for a special-purpose application that was only interesting by itself for a relatively short time, but ended up influencing LZNA, Kraken, Mermaid and Leviathan
Oodle 2.9.10b was released earlier this week.
Data: Mermaid Optimal1 and higher levels compress much faster (>2x is typical in our tests)
- Data: Selkie, Kraken, Leviathan Optimal1+ also compress faster, but less drastically so
Disclaimer on the Mermaid speedup: the ~2x speedup is for 256k chunked encoding in a throughput-bound scenario (e.g. going wide on many chunk encodes at once, our typical use case).
Results will vary with larger chunks or no chunking (more bottlenecked on match finding) or