You have reached the Python page of PDL.
The following Python3 programs may be useful, if handled with care!
LSQFIT and CONTACT require BioPython to be installed, a simple procedure using pip.
See the homepage -- https://pypi.org/project/biopython/
PYMOL_PLANES requires (drum-roll) PYMOL.
LaTeX users may find check_tex a handy way of testing that figures and tables in a tex document are cited in order.
JANUS is a program for generating DNA sequences that encode a given protein.
RAMA
rama will produce a clickable Ramachandran plot within PYMOL.
In the PYMOL terminal, type "run rama.py". This will make the program accessible.
For any PYMOL object (such as a selection) called my_object, type "rama my_object".
This will open a matplotlib window with a Ramachandran plot if the object has
protein residues within it. The plot uses PYMOL's own Ramachandran angles.
The plot points are clickable, and clicking on any of them will create a
temporary object in PYMOL of that residue, centre the view on it, and display sticks.
Clicking within the diagram but not on a plotted point will display the legend
instead of the Ramachandran angles of the last clicked point.
Since the program uses matplotlib, the figure is downloadable as a PNG file.
The plot title includes the name of the PYMOL object in question.
The python file is found here:
rama.py
LSQFIT
lsqfit will match atoms in two files, a reference structure and a moving structure.
If requested, the least-squares best fit of the two sets of coordinates is calculated,
and the moving coordinates are returned as a new PDB file. The rotation-translation operation
is displayed, together with the rotation angle and some CGO arrows for display in PYMOL,
if the screw translation is above 0.2 Angstroms.
Atom selections are easily achieved using a small control file which can be
called simply "fit.in" or "control.txt" or whatever you want.
Lines in fit.in starting with # are ignored, and can be used as comments.
One atom card defines the atom names used.
Multiple lines define the moving (mol1) and fixed (mol2) residues.
mol1 and mol2 lines can be given in any order, but beware as atoms are loaded into
the moving and fixed groups according to the order of the mol2 cards and the mol1 cards.
It is safest to match mol1 and mol2 cards in pairs to ensure the atoms are paired correctly.
mol1 a10 a30
will add residues 10 to 30 of the A chain to the moving group.
One refi card gives a single parameter (Y or N) to indicate if refinement is required.
If not, then the root-mean-square distance is given of the molecules in place, and no PDB is output.
Any atom types can be chosen, allowing models to be superposed using prosthetic groups such
as FAD or haem, or small molecules such as drugs.
You may need to convert HETATM records to ATOM to work with non-protein atoms.
This can be done simply with sed, putting two spaces after ATOM:
sed -i 's\HETATM\ATOM \g' mymodel.pdb
Note that the program is NOT written for speed (it actually carries out the overlap twice, once
using numpy calls and once with a higher function call to BioPython)
but in most cases will give output essentially instantaneously.
The program can be invoked from a UNIX terminal like this:
python3 lsqfit.py -f fixed.pdb -m moving.pdb -o RTout.pdb -c control.txt
The actual python file is found here:
lsqfit.py
A sample control file (which I generally call fit.in) can be found with this link:
fit.in
CONTACT
This works in a similar way to lsqfit, but takes only a single PDB file.
A control file is used to define different groups of atoms, grp1 and grp2.
The program will report any atom in grp1 that is within a chosen distance of any atom
in grp2. Atoms can be chosen by name or element.
Lines beginning with grp1 or grp2 are used to define the residue limits of selections.
grp1 a10 a30
will add residues 10 to 30 (inclusive) from chain A to group 1.
Alternatively a line such as
grp1 a10-30
will do the same job.
Make sure to define both group 1 and group 2.
An element card is used to select atoms to use. An asterisk indicates all element types.
elem O N
will cause the program to report only distances for atom pairs where both atoms are
either oxygen or nitrogen.
dist 3.5
will set the maximum interatomic distance at 3.5 Angstroms.
The results are printed to the terminal, and a LaTeX table file is written to a file,
which can be dropped directly into a LaTeX editor such as Overleaf.
A PyMOL script is also output to draw the contacts.
The program can be invoked from a UNIX terminal like this:
python3 contact.py -f myinputPDB.pdb -c contact.in
The python file is found here:
contact.py
A sample contact.in file can be found here with this link:
contact.in.
PYMOL_PLANES
This file can be run from within PYMOL to calculate the distance of an atom from the
least-squares best-fit plane defined by a group of atoms.
In the PYMOL command line, type
run pymol_plane.py
This will make a command called distance_to_plane available.
Define a plane with 3 or more atoms by atom picking in PYMOL, and renaming the selection.
Then you can type:
distance_to_plane('my_atom', 'myplane')
to find the distance of the atom from the LSQ plane.
The python file is found here:
pymol_plane.py
CHECK_TEX
Fire up python3 and then import the module with
from check_tex import *
If your tex file is called test.tex, then try
check_tex("test.tex")
A log file is output that can be used to help sorting floats into order.
The python file is found here:
check_tex.py
JANUS
This script is designed to produce DNA sequences for high-level protein expression.
The codon-usage table for E. coli is included, but simple modification will allow the program
to work with any chose expression system.
Simple switches on the command line are used to find the protein sequence file (single letter code)
and the filename for the output DNA sequence. Rarer codons may be excluded more or less stringently
with a limit parameter.
Usage:
python JANUS.py -i input.fasta -o output.dna -lim 0.07 -max 250
The python file is found here:
JANUS.py