How to extract the first link from post content (PHP snippet)

How to extract the first link from post content (PHP snippet)

Last verified: September 22, 2026
7 min read
Tutorial
Full-stack developer

Link-roundup themes, press-clipping posts, and “read the original” CTAs all need the first URL inside post HTML. The same scan can promote the first <img> to a featured image when an editor forgets to set one. In both cases you parse post_content (or filtered the_content) and pull attributes from the DOM.

Regex feels faster to write. It fails on real WordPress HTML. Prefer PHP’s DOMDocument, suppress libxml noise, escape every URL you print, and cache results when you run this on archives.

For custom theme and plugin work: WordPress developer.

function wppoland_get_first_link_url( $content ) {
	if ( empty( $content ) ) {
		return false;
	}

	$doc = new DOMDocument();
	libxml_use_internal_errors( true );

	$wrapped = '<?xml encoding="utf-8" ?>' . $content;
	$doc->loadHTML( $wrapped, LIBXML_HTML_NOIMPLIED | LIBXML_HTML_NODEFDTD );

	libxml_clear_errors();

	$links = $doc->getElementsByTagName( 'a' );
	if ( $links->length > 0 ) {
		$href = $links->item( 0 )->getAttribute( 'href' );
		return $href !== '' ? $href : false;
	}

	return false;
}

LIBXML_HTML_NOIMPLIED is optional and behaves differently across libxml versions. If your host’s libxml is old and fragments break, drop those flags and keep the UTF-8 wrapper plus libxml_use_internal_errors( true ).

#Usage in the Loop

$link = wppoland_get_first_link_url( get_the_content( null, false ) );

if ( $link ) {
	printf(
		'<a href="%s" class="read-more-external" rel="noopener noreferrer">Read original</a>',
		esc_url( $link )
	);
}

Use get_post()->post_content when you need raw blocks without the_content filters (shortcodes, embeds). Use get_the_content() when you want the same HTML the front end would render.

#Method 2: first image URL

function wppoland_get_first_image_url( $content ) {
	if ( empty( $content ) ) {
		return false;
	}

	$doc = new DOMDocument();
	libxml_use_internal_errors( true );
	$doc->loadHTML( '<?xml encoding="utf-8" ?>' . $content );
	libxml_clear_errors();

	$images = $doc->getElementsByTagName( 'img' );
	if ( $images->length > 0 ) {
		$src = $images->item( 0 )->getAttribute( 'src' );
		return $src !== '' ? $src : false;
	}

	return false;
}
add_action( 'save_post_post', function( $post_id ) {
	if ( wp_is_post_revision( $post_id ) || wp_is_post_autosave( $post_id ) ) {
		return;
	}
	if ( has_post_thumbnail( $post_id ) ) {
		return;
	}

	$post = get_post( $post_id );
	if ( ! $post ) {
		return;
	}

	$image_url = wppoland_get_first_image_url( $post->post_content );
	if ( ! $image_url ) {
		return;
	}

	$attachment_id = attachment_url_to_postid( $image_url );
	if ( $attachment_id ) {
		set_post_thumbnail( $post_id, $attachment_id );
	}
}, 20 );

attachment_url_to_postid() only resolves media already in the library. External hotlinked images will not become thumbnails without a download step - skip that unless you own the rights and have a clear import path.

#Method 3: regex (reference only)

function wppoland_get_first_link_url_regex( $content ) {
	if ( empty( $content ) ) {
		return false;
	}

	if ( preg_match( '/<a\s[^>]*href=["\']([^"\']+)["\'][^>]*>/i', $content, $matches ) ) {
		return $matches[1];
	}

	return false;
}

Why not ship this in production:

  • Attribute order and single vs double quotes still surprise people.
  • Nested markup and Gutenberg figure/link wrappers confuse naive patterns.
  • Harder to extend to title, rel, or “first external only”.

Keep regex for one-off CLI cleanups, not theme templates.

#Advanced: full data from the first anchor

function wppoland_get_first_link_data( $content ) {
	if ( empty( $content ) ) {
		return false;
	}

	$doc = new DOMDocument();
	libxml_use_internal_errors( true );
	$doc->loadHTML( '<?xml encoding="utf-8" ?>' . $content );
	libxml_clear_errors();

	$links = $doc->getElementsByTagName( 'a' );
	if ( 0 === $links->length ) {
		return false;
	}

	$node = $links->item( 0 );

	return array(
		'url'    => $node->getAttribute( 'href' ),
		'text'   => $node->textContent,
		'title'  => $node->getAttribute( 'title' ),
		'target' => $node->getAttribute( 'target' ),
		'rel'    => $node->getAttribute( 'rel' ),
		'class'  => $node->getAttribute( 'class' ),
	);
}
$data = wppoland_get_first_link_data( get_the_content( null, false ) );
if ( $data && $data['url'] ) {
	printf(
		'<a href="%s" title="%s" target="%s" rel="%s">%s</a>',
		esc_url( $data['url'] ),
		esc_attr( $data['title'] ),
		esc_attr( $data['target'] ? $data['target'] : '_blank' ),
		esc_attr( $data['rel'] ? $data['rel'] : 'noopener noreferrer' ),
		esc_html( $data['text'] )
	);
}

Roundup posts often start with an internal “jump to comments” or related-post link. Skip those:

function wppoland_get_first_external_link( $content ) {
	if ( empty( $content ) ) {
		return false;
	}

	$doc = new DOMDocument();
	libxml_use_internal_errors( true );
	$doc->loadHTML( '<?xml encoding="utf-8" ?>' . $content );
	libxml_clear_errors();

	$site_host = wp_parse_url( home_url(), PHP_URL_HOST );
	$links     = $doc->getElementsByTagName( 'a' );

	foreach ( $links as $link ) {
		$href = trim( $link->getAttribute( 'href' ) );
		if ( '' === $href || '#' === $href[0] ) {
			continue;
		}

		$host = wp_parse_url( $href, PHP_URL_HOST );
		if ( $host && $host !== $site_host ) {
			return $href;
		}
	}

	return false;
}

#Caching so archives stay cheap

Parsing HTML for every card on a category archive is wasteful. Store on save:

function wppoland_cache_first_link( $post_id ) {
	if ( wp_is_post_revision( $post_id ) || wp_is_post_autosave( $post_id ) ) {
		return;
	}

	$post = get_post( $post_id );
	if ( ! $post || 'post' !== $post->post_type ) {
		return;
	}

	$url = wppoland_get_first_link_url( $post->post_content );
	update_post_meta( $post_id, '_wppoland_first_link', $url ? $url : '' );
}
add_action( 'save_post_post', 'wppoland_cache_first_link', 30 );

function wppoland_get_cached_first_link( $post_id = null ) {
	$post_id = $post_id ?: get_the_ID();
	$cached  = get_post_meta( $post_id, '_wppoland_first_link', true );

	if ( '' === $cached ) {
		return false;
	}
	if ( $cached ) {
		return $cached;
	}

	// Backfill for old posts.
	$post = get_post( $post_id );
	$url  = $post ? wppoland_get_first_link_url( $post->post_content ) : false;
	update_post_meta( $post_id, '_wppoland_first_link', $url ? $url : '' );
	return $url;
}

Transients work too; post meta survives cache flushes and is easier to invalidate on save_post.

#Security and edge cases

  • Always esc_url() on output. Never trust content authors blindly.
  • Empty href, javascript: URLs, and data: URIs should be rejected if you only want http(s).
  • Block markup may wrap links in <figure> or use relative paths - DOMDocument still finds the <a>; resolve relative URLs with wp_make_link_relative / site URL as needed.
  • Run extraction in admin or on save when possible so front-end PHP stays thin.

#Gutenberg, classic content, and filtered HTML

post_content in the database stores blocks as HTML comments plus inner HTML. DOMDocument still finds <a> and <img> nodes inside those blocks. If you run apply_filters( 'the_content', $post->post_content ) first, shortcodes and embeds expand - which can insert extra links (share buttons, related posts injected by a plugin). For “first editorial link” semantics, parse raw post_content. For “first link the visitor would see”, parse after the_content.

Classic editor posts and imported XML often contain unclosed tags. That is exactly where regex fails and DOMDocument + libxml_use_internal_errors( true ) earns its keep.

#Reject dangerous schemes

Before you echo a URL as href, allow only http(s) (and optionally relative paths you resolve yourself):

function wppoland_safe_http_url( $url ) {
	$url = trim( $url );
	if ( '' === $url ) {
		return false;
	}
	$parsed = wp_parse_url( $url );
	if ( empty( $parsed['scheme'] ) ) {
		// Site-relative path.
		return esc_url_raw( home_url( $url ) );
	}
	$scheme = strtolower( $parsed['scheme'] );
	if ( ! in_array( $scheme, array( 'http', 'https' ), true ) ) {
		return false;
	}
	return esc_url_raw( $url );
}

Drop javascript:, data:, and vbscript: early. Editors paste weird things; treat content as untrusted input.

#Unit-test the helper

A minimal PHPUnit / WP integration assertion saves regressions when someone “improves” the regex later:

public function test_first_link_skips_hash_and_finds_external() {
	$html = '<p><a href="#comments">Comments</a> <a href="https://example.com/x">X</a></p>';
	$this->assertSame(
		'https://example.com/x',
		wppoland_get_first_external_link( $html )
	);
}

Fixture a post with UTF-8 link text (Norwegian ø, Polish ł) and confirm textContent is not mojibake after loadHTML.

#CLI backfill for old posts

After deploying post-meta caching, backfill once:

wp post list --post_type=post --format=ids | xargs -n1 -I% wp eval 'wppoland_cache_first_link( % );'

Or a small WP_CLI command that batches by ID range so you do not load the entire table into memory on large sites.

#Summary

  1. Prefer DOMDocument over regex for first-link and first-image extraction.
  2. Suppress libxml errors; preserve UTF-8 when loading HTML.
  3. Escape every attribute you print; add rel="noopener noreferrer" for _blank.
  4. Cache on save_post for archive performance.
  5. Filter for external hosts when building aggregator-style CTAs.

Need a production-safe helper packaged as a must-use plugin? Contact with the theme type (classic vs block) and whether you need first link, first image, or both.

Next step

Turn the article into an actual implementation

This block strengthens internal linking and gives readers the most relevant next move instead of leaving them at a dead end.

Want this implemented on your site?

If you want to convert the article into a working site improvement, redesign, or build plan, I can define the scope and implement it.

Related cluster

Explore other WordPress services and knowledge base

Strengthen your business with professional technical support in key areas of the WordPress ecosystem.

Why use DOMDocument instead of regex to extract links?#
DOMDocument understands HTML structure. Regex breaks on attribute order, nested tags, and malformed markup that Gutenberg and classic editors still produce.
How do I handle content with UTF-8 characters?#
Prefer loading with an explicit UTF-8 meta charset wrapper, or convert carefully before loadHTML. Always test with non-ASCII titles and alt text.
Can I extract the first image instead of the first link?#
Yes. Use getElementsByTagName('img') and getAttribute('src') instead of anchors and href.
Should I run this on every page view?#
Cache the result in a transient or post meta on save_post. Parsing HTML on every archive row is unnecessary CPU work.
How do I skip internal links and take the first external URL?#
Loop the DOMNodeList, skip empty hrefs and hash anchors, then compare the host against home_url() before returning.

Need an FAQ tailored to your industry and market? We can build one aligned with your business goals.

Let’s discuss

Related Articles