Imatinib–ABL1 Binding Mode Analysis with ESMFold2
Follow a traceable Imatinib–ABL1 binding mode analysis from sequence and SMILES through three ESMFold2 candidates, confidence ranking, and PoseEdit contacts.
This case study follows one completed Imatinib–ABL1 binding mode analysis from a 271-residue ABL1 kinase-domain sequence and an Imatinib SMILES string to three ESMFold2 complex candidates, confidence ranking, structure conversion, and protein–ligand interaction diagrams. The scores and contacts below describe this prediction run. They do not replace an experimental complex structure, affinity measurement, or biological assay.
The run produced a reviewable answer rather than a single structure file: the original request, three candidate complexes, per-candidate confidence metrics, PAE and distogram data, converted PDB files, interaction summaries, and visual reports remained in one project.

By Mira · Published: July 30, 2026
Evidence note: every run-specific number in this article was checked against the task request, candidate metrics, reports, and interaction files. External references are used only to explain the tools and residue numbering. No RMSD or pose comparison with an experimental ABL1–Imatinib structure was performed in this run.
The question and confirmed input
The task used the phrase “post-docking,” but the actual model input was a protein sequence plus a ligand SMILES, not a pre-docked pose. It asked for an ABL1–Imatinib complex prediction followed by binding-mode analysis of the candidates. Before compute started, the workflow exposed the model and candidate count for confirmation.
| Input | Value used in the run |
|---|---|
| Protein | ABL1 kinase-domain fragment, chain A, 271 amino acids |
| Ligand | Imatinib, chain L |
| Ligand SMILES | CN1CCN(CC1)Cc1ccc(cc1)C(=O)Nc1ccc(C)c(c1)Nc1nccc(n1)c1cccnc1 |
| Model | esmfold2-fast |
| Candidates | 3 |
| Covalent bonds | None specified |

That confirmation step matters. It keeps model choice and sampling scope visible before a relatively expensive run begins, and it makes the result easier to reproduce.
Step 1: submit an all-atom complex request
The workflow used the all-atom input path: one protein sequence and one ligand SMILES in the same request. The final request recorded three recycling loops, 50 sampling steps, three diffusion samples, seed 0, and CIF output. It also requested PAE, distogram, and pair-chain ipTM outputs; embeddings were disabled.
ESMFold2 is a structure-prediction model. In this workflow it generated hypotheses for the complete protein–ligand complex. It did not calculate binding free energy, inhibition potency, residence time, or a biochemical endpoint.
Step 2: rank three candidates without hiding the alternatives
All three candidates completed. The project ranked them by the recorded confidence metrics:
| Rank | Candidate | pLDDT | pTM | ipTM |
|---|---|---|---|---|
| 1 | sample_2 | 92.89 | 0.9612 | 0.9766 |
| 2 | sample_1 | 92.61 | 0.9603 | 0.9766 |
| 3 | sample_0 | 92.85 | 0.9609 | 0.9761 |
Sample 2 ranked first. Its pair-chain ipTM matrix was:
[[0.8412, 0.8733],
[0.7160, 0.8881]]
The run report interpreted the A→L entry (0.8733) as moderate-to-high model confidence. The reverse L→A entry was 0.7160, so this asymmetric matrix should not be reduced to a single interface score. Both values are internal confidence signals, not evidence that the pose is experimentally correct. The narrow spread across the three pTM and ipTM values is useful because it shows that candidate selection was not driven by one isolated score.
Step 3: bridge the format required by the interaction service
The prediction service returned CIF structures, while the downstream binding-mode workflow expected PDB input with a uniquely selectable ligand. The task therefore:
- converted each CIF candidate to PDB with Open Babel;
- renamed the ligand chain from L to B during conversion;
- corrected ligand records from
ATOMtoHETATM; - verified a single ligand selector,
LIG_B_1, containing 37 heavy atoms.
These details are not cosmetic. A correct coordinate file can still fail downstream if the ligand is indistinguishable from protein atoms, the chain identifier is ambiguous, or the residue selector does not match the service contract. Keeping the converted structures beside the originals makes this handoff auditable.
Step 4: calculate and compare interaction diagrams
The workflow sent all three converted candidates to ProteinsPlus, using PoseEdit for interaction diagrams and Protoss for protonation-aware structure preparation.
Each candidate returned the same interaction-category counts:
| Interaction category | sample_0 | sample_1 | sample_2 |
|---|---|---|---|
| Hydrogen bonds | 4 | 4 | 4 |
| Hydrophobic contacts | 3 | 3 | 3 |
| Pi–pi interactions | 2 | 2 | 2 |
| Salt bridges | 0 | 0 | 0 |
| Metal interactions | 0 | 0 | 0 |
| Atom-pair interactions | 0 | 0 | 0 |
| Cation–pi interactions | 0 | 0 | 0 |
The totals in all seven categories were identical across the candidates. Counts alone do not establish that their atom coordinates, geometries, or poses were identical.
For the top-ranked sample 2, the detailed hydrogen-bond records assigned four hydrogen bonds to local residues Glu58, Thr87, Met90, and Ile132. The diagram also assigned pi-stacking contacts to Phe89 and Phe154. PoseEdit separately labeled 11 surrounding residues in its 2D scene.
Local numbering versus canonical ABL1 numbering
The supplied fragment begins at residue 229 of the human ABL1 sequence in UniProt P00519. Local coordinates therefore map to canonical numbering with an offset of 228:
| Contact in the run | Local residue | Canonical ABL1 residue |
|---|---|---|
| Hydrogen bond | Glu58 | Glu286 |
| Hydrogen bond | Thr87 | Thr315 |
| Pi stacking | Phe89 | Phe317 |
| Hydrogen bond | Met90 | Met318 |
| Hydrogen bond | Ile132 | Ile360 |
| Pi stacking | Phe154 | Phe382 |
This mapping prevents a common reporting error: comparing a fragment-local label such as Thr87 directly with full-length ABL1 literature. It aligns identifiers only. It does not prove that the predicted geometry matches an experimental pose.
What the result files make reviewable
The final 69 MB tar.gz archive contained 89 files. The main groups were:
- three raw CIF candidates and their metrics;
- PAE, distogram, and pair-chain ipTM arrays;
- converted PDB structures and a best-candidate file;
- per-candidate PoseEdit and Protoss inputs and outputs;
- interaction counts, 2D diagrams, reports, and board data.
The large distogram arrays accounted for most of the archive size. More important than the count is provenance: a reviewer can move from the summary table back to the candidate structure, raw metric, or interaction record that produced it.
What this run supports, and what it does not
This workflow supports three practical uses:
- generating several protein–ligand pose hypotheses from a stated sequence and SMILES;
- ranking those hypotheses with recorded model-confidence outputs;
- organizing candidate-specific contact diagrams for expert review and follow-up.
It does not establish binding affinity, inhibition, selectivity, cellular activity, or clinical relevance. It also does not validate the predicted pose against crystallography. A human ABL1–Imatinib X-ray structure is available as PDB 2HYY, but this task did not perform a structural alignment, ligand RMSD calculation, or residue-by-residue comparison to that reference.
A defensible next step would define a protonation and preparation protocol, align the predicted complex to an experimental reference, calculate protein and ligand RMSD, inspect contact conservation, and then decide whether molecular dynamics or an affinity-oriented method is warranted. Those are new calculations, not conclusions that can be inferred from the current images.
A reusable protein–ligand review checklist
Before promoting any predicted complex to a downstream study, check:
- Identity: sequence boundaries, ligand structure, protonation state, chain IDs, and covalent bonds.
- Sampling: model, seed, loop count, sampling steps, and number of candidates.
- Confidence: per-candidate pLDDT, pTM, ipTM, PAE, and the spread between candidates.
- Format handoff: atom records, residue names, chain IDs, and the unique ligand selector.
- Numbering: fragment-local residue numbers mapped explicitly to the reference sequence.
- Validation boundary: model confidence separated from comparison with an experimental structure or assay.
For a different example of how Mira preserves inputs, methods, outputs, and limits around a computational chemistry task, see the phenylethyl resorcinol orbital and ESP workflow.
To run a protein–ligand question in the same project context, start a Mira project.
FAQ
Why did the workflow generate three ESMFold2 candidates?
Multiple candidates expose sampling variation and make it possible to compare confidence and interaction profiles instead of treating the first returned structure as definitive. This run requested three diffusion samples.
What do pLDDT, pTM, and ipTM prove?
They report different aspects of model confidence in the predicted structure and interfaces. High values can help rank candidates, but they do not prove that a ligand pose is experimentally correct or that the ligand binds with a particular affinity.
Why convert CIF to PDB and change the ligand to HETATM records?
The downstream interaction service expected a PDB structure with a uniquely identifiable ligand. Converting the files, assigning chain B, and using HETATM records allowed the task to verify the selector LIG_B_1 before analysis.
Do the predicted contacts validate Imatinib binding to ABL1?
No. They are contacts in the predicted candidates. This run did not compare the pose with PDB 2HYY or perform an affinity or biological assay, so the contacts should be treated as hypotheses for review.

