# Prevent Googlebot-Video from crawling - robots.txt / X-Robots-Tag?

**URL:** <https://community.wowza.com/t/prevent-googlebot-video-from-crawling-robots-txt-x-robots-tag/53723>\
**Category:** Wowza Streaming Engine\
**Created:** [July 29, 2019, 11:55am UTC](https://community.wowza.com/t/prevent-googlebot-video-from-crawling-robots-txt-x-robots-tag/53723 "2019-07-29T11:55:41Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![Elio\_Wahlen](https://avatars.discourse-cdn.com/v4/letter/e/3be4f8/32.png) [@Elio\_Wahlen](https://community.wowza.com/u/Elio_Wahlen)\
**Post date:** [July 29, 2019, 11:55am UTC](https://community.wowza.com/t/prevent-googlebot-video-from-crawling-robots-txt-x-robots-tag/53723/1 "2019-07-29T11:55:41Z")

</div>

Our VOD-Servers are being crawled by the Googlebot (User Agent “Googlebot-Video/1.0”)

We have many terabytes of data there and we want to avoid that traffic.

I would like to put a robots.txt to prevent the googlebot from continuing.  
Wowza runs on its own subdomain so the robots.txt would need to be hosted from Wowza itself.  
Is there a good practice for that

Or does wowza support ways of setting up the X-Robots-Tag in the HTTP headers?

---

<div class="post-metadata">

**Author:** ![Rose\_Power-Wowza\_Com](https://sea2.discourse-cdn.com/flex002/user_avatar/community.wowza.com/rose_power-wowza_com/32/431_2.png) [@Rose\_Power-Wowza\_Com](https://community.wowza.com/u/Rose_Power-Wowza_Com)\
**Post date:** [July 29, 2019, 7:07pm UTC](https://community.wowza.com/t/prevent-googlebot-video-from-crawling-robots-txt-x-robots-tag/53723/2 "2019-07-29T19:07:03Z")

</div>

Hi @Elio Wahlen robots.txt is added to a web server; WSE is not really a web server though, but you can add custom headers to your HLS chunklist though.

[https://www.wowza.com/docs/how-to-add-custom-playlist-headers-to-apple-hls-manifests](https://www.wowza.com/docs/how-to-add-custom-playlist-headers-to-apple-hls-manifests)

---

<div class="post-metadata">

**Author:** ![Elio\_Wahlen](https://avatars.discourse-cdn.com/v4/letter/e/3be4f8/32.png) [@Elio\_Wahlen](https://community.wowza.com/u/Elio_Wahlen)\
**Post date:** [July 29, 2019, 8:58pm UTC](https://community.wowza.com/t/prevent-googlebot-video-from-crawling-robots-txt-x-robots-tag/53723/3 "2019-07-29T20:58:40Z")

</div>

Thanks for this good and helpful answer.

Is there also a way for MPEG-DASH and HDS (i know HDS is dying, but we still need to support it for a while for our customers)?

I just saw in our logs that google was using the HDS manifest to access the vod-streams…

---

<div class="post-metadata">

**Author:** ![Davi\_Jones](https://avatars.discourse-cdn.com/v4/letter/d/f6c823/32.png) [@Davi\_Jones](https://community.wowza.com/u/Davi_Jones)\
**Post date:** [July 30, 2019, 1:58am UTC](https://community.wowza.com/t/prevent-googlebot-video-from-crawling-robots-txt-x-robots-tag/53723/4 "2019-07-30T01:58:26Z")

</div>

Thanks. I had the some issue. Help me a lot . 🙂 [@arquiteto](https://www.tudopraobra.com.br/empresas/arquitetura-de-interiores)

---

<div class="post-metadata">

**Author:** ![Elio\_Wahlen](https://avatars.discourse-cdn.com/v4/letter/e/3be4f8/32.png) [@Elio\_Wahlen](https://community.wowza.com/u/Elio_Wahlen)\
**Post date:** [July 30, 2019, 3:48pm UTC](https://community.wowza.com/t/prevent-googlebot-video-from-crawling-robots-txt-x-robots-tag/53723/5 "2019-07-30T15:48:24Z")

</div>

Dear [@Rose Power-Wowza Community Manager](http://community.wowza.com/community/questions/52460/prevent-googlebot-video-from-crawling-robotstxt.html#)

Unfortunately I now realize that you got me wrong.  
The chunklist headers are not equivalent to the HTTP headers that I meant. Specifically I am talking about the X-Robots-Tag HTTP header that seem to be standard.  
Please see here: [https://developers.google.com/search/reference/robots\_meta\_tag](https://developers.google.com/search/reference/robots_meta_tag)

Is there any way for sending custom HTTP response headers across all HTTP based connections?  
If not - what do you recommend? Is reverse-proxying e.g. with nginx a way to go? How would that look like?

I can imagine there are a lot of users that don’t want googlebot and possibly other crawlers to download all their video content because of different reasons (traffic, performance, legal circumstances, etc). It would be wise to have a plan here.

Best wishes, Elio

---

<div class="post-metadata">

**Author:** ![Rose\_Power-Wowza\_Com](https://sea2.discourse-cdn.com/flex002/user_avatar/community.wowza.com/rose_power-wowza_com/32/431_2.png) [@Rose\_Power-Wowza\_Com](https://community.wowza.com/u/Rose_Power-Wowza_Com)\
**Post date:** [July 31, 2019, 3:22pm UTC](https://community.wowza.com/t/prevent-googlebot-video-from-crawling-robots-txt-x-robots-tag/53723/6 "2019-07-31T15:22:08Z")

</div>

Apologies @Elio Wahlen. You can add custom http headers to hls/dash/hds by adding

`httpUserHTTPHeaders` property to your application (this is similar to the Access-Control-Allow-Origin cors http headers).

Here’s an example of how this can be added:  
[https://www.wowza.com/docs/how-to-stream-from-an-android-device-to-the-google-chromecast-device](https://www.wowza.com/docs/how-to-stream-from-an-android-device-to-the-google-chromecast-device)

* * *

the value for the property in your case would be:

```auto
X-Robots-Tag: noarchive

```

etc. It’s a pipe-delimited list, so you can add multiple headers.

if you prefer to host a robots.txt file, then you would need to have a custom HTTPProvider that handles requests for robots.txt; this is similar to how WSE handles  
http://:1935/crossdomain.xml

[https://www.wowza.com/docs/how-to-create-an-http-provider](https://www.wowza.com/docs/how-to-create-an-http-provider)

---

<div class="post-metadata">

**Author:** ![Elio\_Wahlen](https://avatars.discourse-cdn.com/v4/letter/e/3be4f8/32.png) [@Elio\_Wahlen](https://community.wowza.com/u/Elio_Wahlen)\
**Post date:** [August 4, 2019, 12:22am UTC](https://community.wowza.com/t/prevent-googlebot-video-from-crawling-robots-txt-x-robots-tag/53723/7 "2019-08-04T00:22:00Z")

</div>

@Rose Power-Wowza Community Manager thanks very much.

httpUserHTTPHeaders works nicely!

I think it would help people to have a general tutorial or manual entry about adding these super useful custom http headers. For now it seems to be hidden in two more special tutorials.

All the best, Elio

---

<div class="post-metadata">

**Author:** ![Rose\_Power-Wowza\_Com](https://sea2.discourse-cdn.com/flex002/user_avatar/community.wowza.com/rose_power-wowza_com/32/431_2.png) [@Rose\_Power-Wowza\_Com](https://community.wowza.com/u/Rose_Power-Wowza_Com)\
**Post date:** [August 4, 2019, 7:39pm UTC](https://community.wowza.com/t/prevent-googlebot-video-from-crawling-robots-txt-x-robots-tag/53723/8 "2019-08-04T19:39:21Z")

</div>

Fabulous idea! Thank you so much for the feedback.
