Imatinib–ABL1 Binding Mode Analysis with ESMFold2

Follow a traceable Imatinib–ABL1 binding mode analysis from sequence and SMILES through three ESMFold2 candidates, confidence ranking, and PoseEdit contacts.

Published: Jul 30, 202610 min read

This case study follows one completed Imatinib–ABL1 binding mode analysis from a 271-residue ABL1 kinase-domain sequence and an Imatinib SMILES string to three ESMFold2 complex candidates, confidence ranking, structure conversion, and protein–ligand interaction diagrams. The scores and contacts below describe this prediction run. They do not replace an experimental complex structure, affinity measurement, or biological assay.

The run produced a reviewable answer rather than a single structure file: the original request, three candidate complexes, per-candidate confidence metrics, PAE and distogram data, converted PDB files, interaction summaries, and visual reports remained in one project.

Mira board showing the Imatinib–ABL1 candidate ranking, 2D interaction diagram, and 3D binding structure

By Mira · Published: July 30, 2026

Evidence note: every run-specific number in this article was checked against the task request, candidate metrics, reports, and interaction files. External references are used only to explain the tools and residue numbering. No RMSD or pose comparison with an experimental ABL1–Imatinib structure was performed in this run.


The question and confirmed input

The task used the phrase “post-docking,” but the actual model input was a protein sequence plus a ligand SMILES, not a pre-docked pose. It asked for an ABL1–Imatinib complex prediction followed by binding-mode analysis of the candidates. Before compute started, the workflow exposed the model and candidate count for confirmation.

InputValue used in the run
ProteinABL1 kinase-domain fragment, chain A, 271 amino acids
LigandImatinib, chain L
Ligand SMILESCN1CCN(CC1)Cc1ccc(cc1)C(=O)Nc1ccc(C)c(c1)Nc1nccc(n1)c1cccnc1
Modelesmfold2-fast
Candidates3
Covalent bondsNone specified

Two-dimensional structure generated for the supplied Imatinib SMILES

That confirmation step matters. It keeps model choice and sampling scope visible before a relatively expensive run begins, and it makes the result easier to reproduce.

Step 1: submit an all-atom complex request

The workflow used the all-atom input path: one protein sequence and one ligand SMILES in the same request. The final request recorded three recycling loops, 50 sampling steps, three diffusion samples, seed 0, and CIF output. It also requested PAE, distogram, and pair-chain ipTM outputs; embeddings were disabled.

ESMFold2 is a structure-prediction model. In this workflow it generated hypotheses for the complete protein–ligand complex. It did not calculate binding free energy, inhibition potency, residence time, or a biochemical endpoint.

Step 2: rank three candidates without hiding the alternatives

All three candidates completed. The project ranked them by the recorded confidence metrics:

RankCandidatepLDDTpTMipTM
1sample_292.890.96120.9766
2sample_192.610.96030.9766
3sample_092.850.96090.9761

Sample 2 ranked first. Its pair-chain ipTM matrix was:

[[0.8412, 0.8733],
 [0.7160, 0.8881]]

The run report interpreted the A→L entry (0.8733) as moderate-to-high model confidence. The reverse L→A entry was 0.7160, so this asymmetric matrix should not be reduced to a single interface score. Both values are internal confidence signals, not evidence that the pose is experimentally correct. The narrow spread across the three pTM and ipTM values is useful because it shows that candidate selection was not driven by one isolated score.

Step 3: bridge the format required by the interaction service

The prediction service returned CIF structures, while the downstream binding-mode workflow expected PDB input with a uniquely selectable ligand. The task therefore:

  1. converted each CIF candidate to PDB with Open Babel;
  2. renamed the ligand chain from L to B during conversion;
  3. corrected ligand records from ATOM to HETATM;
  4. verified a single ligand selector, LIG_B_1, containing 37 heavy atoms.

These details are not cosmetic. A correct coordinate file can still fail downstream if the ligand is indistinguishable from protein atoms, the chain identifier is ambiguous, or the residue selector does not match the service contract. Keeping the converted structures beside the originals makes this handoff auditable.

Step 4: calculate and compare interaction diagrams

The workflow sent all three converted candidates to ProteinsPlus, using PoseEdit for interaction diagrams and Protoss for protonation-aware structure preparation.

Each candidate returned the same interaction-category counts:

Interaction categorysample_0sample_1sample_2
Hydrogen bonds444
Hydrophobic contacts333
Pi–pi interactions222
Salt bridges000
Metal interactions000
Atom-pair interactions000
Cation–pi interactions000

The totals in all seven categories were identical across the candidates. Counts alone do not establish that their atom coordinates, geometries, or poses were identical.

For the top-ranked sample 2, the detailed hydrogen-bond records assigned four hydrogen bonds to local residues Glu58, Thr87, Met90, and Ile132. The diagram also assigned pi-stacking contacts to Phe89 and Phe154. PoseEdit separately labeled 11 surrounding residues in its 2D scene.

PoseEdit two-dimensional interaction diagram for the top-ranked sample 2

Local numbering versus canonical ABL1 numbering

The supplied fragment begins at residue 229 of the human ABL1 sequence in UniProt P00519. Local coordinates therefore map to canonical numbering with an offset of 228:

Contact in the runLocal residueCanonical ABL1 residue
Hydrogen bondGlu58Glu286
Hydrogen bondThr87Thr315
Pi stackingPhe89Phe317
Hydrogen bondMet90Met318
Hydrogen bondIle132Ile360
Pi stackingPhe154Phe382

This mapping prevents a common reporting error: comparing a fragment-local label such as Thr87 directly with full-length ABL1 literature. It aligns identifiers only. It does not prove that the predicted geometry matches an experimental pose.

What the result files make reviewable

The final 69 MB tar.gz archive contained 89 files. The main groups were:

  • three raw CIF candidates and their metrics;
  • PAE, distogram, and pair-chain ipTM arrays;
  • converted PDB structures and a best-candidate file;
  • per-candidate PoseEdit and Protoss inputs and outputs;
  • interaction counts, 2D diagrams, reports, and board data.

The large distogram arrays accounted for most of the archive size. More important than the count is provenance: a reviewer can move from the summary table back to the candidate structure, raw metric, or interaction record that produced it.

What this run supports, and what it does not

This workflow supports three practical uses:

  1. generating several protein–ligand pose hypotheses from a stated sequence and SMILES;
  2. ranking those hypotheses with recorded model-confidence outputs;
  3. organizing candidate-specific contact diagrams for expert review and follow-up.

It does not establish binding affinity, inhibition, selectivity, cellular activity, or clinical relevance. It also does not validate the predicted pose against crystallography. A human ABL1–Imatinib X-ray structure is available as PDB 2HYY, but this task did not perform a structural alignment, ligand RMSD calculation, or residue-by-residue comparison to that reference.

A defensible next step would define a protonation and preparation protocol, align the predicted complex to an experimental reference, calculate protein and ligand RMSD, inspect contact conservation, and then decide whether molecular dynamics or an affinity-oriented method is warranted. Those are new calculations, not conclusions that can be inferred from the current images.

A reusable protein–ligand review checklist

Before promoting any predicted complex to a downstream study, check:

  1. Identity: sequence boundaries, ligand structure, protonation state, chain IDs, and covalent bonds.
  2. Sampling: model, seed, loop count, sampling steps, and number of candidates.
  3. Confidence: per-candidate pLDDT, pTM, ipTM, PAE, and the spread between candidates.
  4. Format handoff: atom records, residue names, chain IDs, and the unique ligand selector.
  5. Numbering: fragment-local residue numbers mapped explicitly to the reference sequence.
  6. Validation boundary: model confidence separated from comparison with an experimental structure or assay.

For a different example of how Mira preserves inputs, methods, outputs, and limits around a computational chemistry task, see the phenylethyl resorcinol orbital and ESP workflow.

To run a protein–ligand question in the same project context, start a Mira project.

FAQ

Why did the workflow generate three ESMFold2 candidates?

Multiple candidates expose sampling variation and make it possible to compare confidence and interaction profiles instead of treating the first returned structure as definitive. This run requested three diffusion samples.

What do pLDDT, pTM, and ipTM prove?

They report different aspects of model confidence in the predicted structure and interfaces. High values can help rank candidates, but they do not prove that a ligand pose is experimentally correct or that the ligand binds with a particular affinity.

Why convert CIF to PDB and change the ligand to HETATM records?

The downstream interaction service expected a PDB structure with a uniquely identifiable ligand. Converting the files, assigning chain B, and using HETATM records allowed the task to verify the selector LIG_B_1 before analysis.

Do the predicted contacts validate Imatinib binding to ABL1?

No. They are contacts in the predicted candidates. This run did not compare the pose with PDB 2HYY or perform an affinity or biological assay, so the contacts should be treated as hypotheses for review.