TheUnknownBlog

Back

The phone is a Redmi K30 Ultra, codename cezanne, with a Dimensity 1000+: four Cortex-A77 big cores and four Cortex-A55 little cores. The bootloader was already unlocked. A while ago my mom told me the phone had started rebooting endlessly while charging and then got stuck on the Mi logo. I didn’t think much of it at the time — I assumed it was an automatic system update that had failed halfway through (this phone’s MIUI version is quite old; if I remember correctly, it doesn’t even have A/B partitions).

I figured a fresh flash would fix it. Instead, the whole thing spiraled into ARM error syndrome registers, device trees, and MediaTek CPU hotplug. The system was saved — at the cost of the four big cores. So yes, when the four big cores were in trouble, the four little cores really did just stand by and watch. More precisely: the four big cores were left offline, and MIUI booted on the four Cortex-A55 little cores alone. No replacement motherboard, no replacement battery, no heat gun on the CPU. As for whether the CPU itself is actually broken — probably yes, but it’s hard to say exactly what went wrong.

This debugging session was a joint effort between me and Codex. I pressed buttons, watched the screen, and switched boot modes; it read logs, dug through source code, wrote tools, and modified images on the Mac. Several times we thought “this is it,” only for the phone to fall back into a boot loop moments later. All I can say is: praise Lord Astra’s divine power.

Earlier, I had flashed the official Fastboot ROM on another Windows machine with MiFlash, choosing “clean all” — MIUI 13 based on Android 12. The flash completed without a single error; every partition returned OKAY. Yet the phone still restarted around the REDMI logo, never even reaching the MIUI boot animation. Fastboot, on the other hand, was rock solid — it flashed fine and reported no errors.

GPT, that idiot model, has a bad habit of never taking what users say at face value. The first time I asked it, instead of questioning what was wrong with the phone, it questioned what was wrong with me 😅. It made me verify whether it was “really stuck on the Mi logo” and whether I had “downloaded the right ROM package.” It infuriated me 😡. So I started a new session and emphasized from the very beginning: the ROM was correct, the flash had genuinely succeeded, and don’t ask me to check the model number again. I had no known-good battery, no bench power supply, and no idea how the phone had originally broken. All I had was this Mac, a USB cable, and a Fastboot that was still alive. The rest was on it to figure out.

You might also be wondering: how many tokens did I burn? I got everything done within a single session. At current API prices, that came to $58.19 — about 60% of the weekly quota on my $20 account. Put that way, it’s like I solved the whole thing (at least partially) for the equivalent of about $3 — a pretty good deal 😁.

Fastboot Works Fine, So Why Won’t the System Boot?

If the CPU were truly broken, how could Fastboot survive my flashing without any issues? But thinking it over carefully, and recalling what I wrote in Learning Linux Kernel (Part 2) - Bootstrapping, breaking the boot process into stages makes it clear. Fastboot runs inside the bootloader — the code it executes, the hardware it enables, and its power-management state are all extremely limited. The deeper CPU initialization, frequency scaling, and driver loading only happen once the kernel starts; Android userspace comes even later.

Of course, the reverse doesn’t hold either: a system that won’t boot doesn’t mean the flash failed. Later, once we had a root shell in Recovery, we read back the entire 64 MiB boot partition and computed its SHA-256 — it matched the stock image exactly. The 128 MiB Recovery partition also matched the diagnostic image flashed at the time. So at least the UFS was ruled out. Great — one fewer suspect on the list.

Where Do the Logs Come From?

Normally, the first reaction to an Android problem is adb logcat. No such luxury this time — we couldn’t even get into Recovery, so there was no ADB shell to work with.

Fortunately, the MediaTek bootloader left a door open. Following the strings and source code in LK, you can find commands like oem lkmsg, oem lpmsg, and oem dump_pllk_log, which let you read out the saved kernel console, message buffers, and preloader/LK logs.

So Lord Codex spent some tokens writing a small tool in Python with PyUSB and libusb, speaking the USB protocol directly: send a command, read INFO, receive by length on DATA, and finally wait for OKAY. We finally fished out some logs: a 256 KiB console buffer, a 64 KiB message buffer, and bootloader logs.

Then came the long cycle of reboot → crash → read logs in the bootloader. As for why we rebooted no fewer than ten times — it’s because the logs we pulled out were frequently byte-for-byte identical to the previous round 😅. LK restores old exception records from expdb. The console buffer is a fixed-size circular buffer: the front half holds new content while the back half may still hold leftovers from before. One time Recovery was visibly running just fine, yet the exported OEM log was byte-identical to the previous crash.

So Lord Codex started numbering each captured log, storing the raw bytes, timestamps, and hashes, and comparing which content had actually changed. The first decent kernel exception landed in the LED/backlight device-registration path. The stack looked roughly like this (with some intermediate frames omitted):

__pi_strcmp
class_find_device
of_led_classdev_register
devm_of_led_classdev_register
mtk_leds_probe
text

Black screen, and the stack happens to contain LED code — sounds plausible, right? We almost started suspecting the backlight hardware. But strcmp receiving a bad pointer only tells you where the program hit the wall. Whether the driver itself passed something wrong, or something earlier had corrupted the data, you can’t tell from this alone. To see more clearly, we left the kernel untouched and only added to the boot arguments:

ignore_loglevel loglevel=8 initcall_debug
text

This time the log showed mtk_leds_init returning 0 normally. The “broken backlight” theory could no longer explain what we were seeing. The really interesting part came later — here are a few key lines:

ecc error,irq_index:32, misc0_el1: 0x0000bd02a0018184, status_el1: 0x6e000007
ecc error,irq_index:36, misc0_el1: 0x0000e00180000e80, status_el1: 0x4e000006
ecc error,irq_index:33, misc0_el1: 0x0000ca0180001e40, status_el1: 0x4e000006
read_timeout_handler:33: read timeout
Insufficient stack space to handle exception!
text

This was far more specific than “some driver dereferenced a null pointer.” The corresponding MediaTek cache_parity driver reads ARM registers like ERXSTATUS_EL1 and ERXMISC0_EL1. In 0x6e000007, both the valid bit and the uncorrected-error flag are set, followed by a bus read timeout.

With the logs pointing toward the CPU/cache path, the most direct controlled experiment was: boot with fewer cores enabled. We added:

maxcpus=1
text

We redid this on the Recovery with detailed logging enabled, and only then got a verifiable result. Recovery finally booted. Next we discovered that Xiaomi’s stock Recovery has no ADB. Under stock Recovery the computer saw unauthorized, and no authorization prompt ever appeared on the phone. I tapped “Connect with MiAssistant,” the device switched to sideload, and a regular adb shell still returned closed. The final workaround: on the Recovery that could already boot, modify the ADB properties and init configuration in the ramdisk — kernel, device tree, and menus all untouched — to give us a temporary root ADB. With a shell, we could finally stop guessing and read logs directly. All in all, after a long debugging process, the final Recovery could run with the four little cores enabled. Very, very good 👍.

online:   0-3
present:  0-3
possible: 0-7
text

Can We Disable Just the Bad Core?

At this point my thinking was naive: shutting down all four big cores felt too wasteful — could we find out exactly which one was broken and disable only that one? The bootloader was already unlocked; might as well use it. So we started building Recovery variants with “four little cores + one big core allowed at a time.”

The actual approach: delete enable-method from the other big-core nodes so they fail CPU operations initialization, adjust cpu-map, and meanwhile keep the original node order and hardware IDs unchanged, so we wouldn’t mix up core numbering during testing.

CPU 4 almost convinced me it was fine: MIDR 0x411fd0d0 confirmed it as a Cortex-A77, frequency 2.6 GHz. Then it restarted by itself — nobody touched it. When ADB reconnected, uptime had already reset to zero and started counting again; then the connection dropped again. After lengthy experiments and many manual interventions to get back into the bootloader, it seemed that enabling any big core led to all kinds of problems…

ConfigurationWhat we actually observed
Four little cores + CPU 4Entered Recovery; three pinned-CPU checksums came back correct, then uptime reset with no human intervention; fresh ECC/timeout records were captured during this configuration
Four little cores + CPU 5First round auto-returned to Fastboot and yielded a stale log; on retry there were new CPU 5 work records and bus timeouts, but no stable ADB
Four little cores + CPU 6Briefly entered Recovery; fresh ECC records were captured afterward, including an uncorrected-error flag
Four little cores + CPU 7A kernel exception in alloc_vmap_area was captured

There was also some fault of my own here: midway through I kept saying “black screen and reboot,” but looking back, a few times I only saw the black screen and never confirmed it had actually gone through the boot sequence again — it might have been stuck, or just booting slowly.

After All That Fiddling

Since per-core testing couldn’t find any keepable big-core combination, the next step was to mask them all off and get the phone usable first. The natural idea: delete enable-method from all four big cores and keep only the four little cores. But this version also stayed on a black screen — and the log gave a reason different from before. More than thirty seconds into boot, the kernel began printing this over and over:

[CPU][EEM]@eem_init01():3389, get_volt(EEM_DET_B) = 0x00000060, VBOOT = 0x00000038
text

MediaTek’s EEM voltage initialization was waiting for the big cluster’s voltage to reach its expected value. The CPUs had been removed from the topology, but the accompanying voltage-initialization flow hadn’t been removed with them. This log contained no fresh ECC records — it was basically just this one line spamming the buffer. That meant this time it was at least stuck on an initialization wait, and the earlier CPU-fault explanation couldn’t cover everything we were seeing.

Fine — the driver has a mind of its own. So we restored the Recovery that had demonstrably booted before, left the device tree alone, and kept using maxcpus=1. It worked again. Online CPUs stayed at 0-3, and cache/bus error status stayed at zero the whole time. Finally, a baseline we could move forward from.

All those earlier rounds had only touched Recovery; the Android boot partition had kept its stock contents the whole time. Only at this step did we write a modified boot image into it for the first time. The final solution changed neither kernel code nor Android’s ramdisk and device tree — it simply carried over the boot arguments from the successful Recovery:

bootopt=64S3,32N2,64N2 buildvariant=user ignore_loglevel loglevel=8 initcall_debug k30u_diag=singlecpu_boot maxcpus=1
text

k30u_diag was a marker we stuffed in ourselves to make the image easy to identify; the one actually doing the work was maxcpus=1. The verbose logging arguments stayed in this time too.

Flash boot, reboot. After about two minutes of waiting, REDMI passed. The MIUI animation appeared. Just let it spin — no button presses, no image swaps. And then it really did enter the system!!! It was all worth it. Lord Astra truly deserves the title of AGI — phone repair folks really are going to be out of a job…

With only four A55s left, every bit of savings counts. I had Codex look at the background services and disable the advertising service, usage analytics, the minus-one screen (App Vault), Quick Apps and its auxiliary component, and the Taplus portal — using pm disable-user --user 0. I also wanted Codex to set animations to 0.5x and turn off blur effects, but ADB complained about missing WRITE_SECURE_SETTINGS. MIUI has a separate “USB debugging (Security settings)” toggle, and turning it on actually requires inserting a SIM card. I just wanted to change an animation scale…

Oh, and we disabled system updates while we were at it. Don’t want an automatic update coming along and overwriting the modified boot arguments.

References

My Old Phone Wouldn't Boot Anymore
https://theunknownth.ing/blog/redmi-k30-ultra-bootloop/en
Author TheUnknownThing
Published at October 2, 2026
Comment seems to stuck. Try to refresh?✨